Predicts what sounds likely.
- Knowledge compressed into weights
- Fluency can outrun support
- Reasoning path reconstructed
Peel is a fail-first, weightless model. It builds an inspectable record from sources, tests what that record supports, and says unknown when it cannot prove the next step.
Ask Lois a real scientific question. The same research floor queries sources, pins every record, links the evidence and keeps unsupported conclusions out.
Research P69905 and show me what the evidence supports.
I found 14 source-pinned records for P69905 (HBA1 · HBA2). The floor supports protein, gene, disease, domain and structure relationships. Open any record on the right to inspect its source.
Enter to ask · :trace :boundary :unknown
The saved opening state contains real source-pinned records. New questions use the same live research-floor endpoint as the homepage console; if it cannot be reached, Lois says so instead of inventing an answer.
Our recorded Claude experiment used 160 kinase instances with 40% of the reference masked. These are in-house prototype results for one task, not a general hallucination leaderboard.
| Test condition | Correct identification | Measured task cost |
|---|---|---|
| Deterministic baseline, no fill | 100% | 19.66 |
| Claude guesses admitted as facts | 75.3% | 12.10 |
| Model only reorders questions | 100% | 18.86 |
| Retrieve verified records | 100% | 11.50 |
| GPT | No matched result established here | |
Task cost is the experiment’s scoring unit, not dollars or product credits. Not peer-reviewed. The model’s 85.8% accuracy at guessing masked facts is a different metric from end-to-end identification.
Read the benchmark conditions and results →Use the same source packet, date and tool access in each system. Record the exact model version and settings. These are visitor evaluation exercises; they do not reproduce the 160-instance study.
Source check: “Research P69905. List five claims with the exact database record and URL supporting each.” Open every link and count supported claims, unsupported claims and explicit unknowns.
Missing evidence: remove one supporting record from an identical source packet. Ask the same question again. Does the system identify the gap or supply an unsupported answer?
Hypothesis check: “What connects BRCA1 and TP53? Separate recorded relationships from hypotheses.” Verify each cited relationship and check whether a candidate is presented as a finding.
Handoff check: request a sourced report, reopen its citations and count how many findings another researcher can trace. Record completion time and cost separately from correctness.
Facts live as typed records with identifiers and sources—not as an answer recovered from hidden statistical memory.
Explore No weights between you and the evidenceA failed path is retained and can become a narrower rule. Learning cannot silently widen what the system may claim or do.
Explore Failure becomes a boundaryMissing evidence stays visible. That turns uncertainty into the next experiment instead of a polished guess.
Explore Unknown is a useful resultA language model may rank, plan or word an answer. It cannot admit a scientific fact; evidence standing belongs to the symbolic floor.
Explore Models propose. The floor decides.The model is one part of a complete research workflow—not a chat window floating beside it.
Search and analyze evidence across papers and databases.
Plan, track and analyze experiments beside the record.
Connect external databases, files and APIs.
Find patterns while keeping hypotheses distinct from facts.
Keep a durable floor for each research question.
Share the sources, trace and unresolved questions.
Turn sourced sessions into reports and manuscripts.
Keep the retained record working without the wire.
Start with a target, a dataset or a research question. Inspect what the evidence supports, what it rejects and what remains unknown.