The Trust Index
Every brief the agent produces carries a Trust Index: a single 0–100 number, a plain-language trust band, and six facet checks that show exactly why the score landed where it did. It is a summary of how well the evidence behind a brief actually holds up — not a rating of how confident the writing sounds.
#How the score works
As the agent assembles a brief, it inspects the evidence supporting each claim and rolls those judgements into one index from 0 to 100. The index is deliberately conservative: a brief earns a high score by resting on reputable, full-text sources whose claims are grounded in cited passages, corroborated across independent studies, and stable over time. Thin sourcing, a single unreplicated result, or a claim the source text does not actually support all pull the number down.
The index is derived from six facet checks. Each facet examines a different way evidence can be strong or weak, and each is marked OK, Caution, Weak, or N/A. The facets are not simply averaged — a facet that speaks directly to a claim’s reliability, such as grounding, weighs more heavily than one that may not apply to the question at hand. Because the score is decomposed this way, the number is never a black box: you can always open the facets to see what earned it.
Each facet mark carries a specific meaning that is consistent across every brief:
- OK — the check passed; the evidence meets the bar for this facet.
- Caution — partially met; there is a soft spot worth a human’s eye.
- Weak — the check failed or the supporting evidence is notably thin.
- N/A — the facet does not apply to this question, so it is excluded from the score rather than penalising it.
#The six facet checks
Together the six facets cover the questions a careful reader would ask before trusting a finding: where did this come from, does the text really say it, how good is the evidence, has anyone reproduced it, does it still hold, and were the tools behind it sound.
#Provenance
Are the sources reputable and traceable? Provenance looks at whether each cited work is a real, identifiable publication from the life-science literature — with a resolvable identifier and a legitimate venue — rather than an untraceable or low-integrity source. Briefs that lean on well-established, full-text papers score well here; those resting on sources that cannot be cleanly traced are flagged.
#Grounding
Is every claim actually tied to a cited passage? Grounding is the check that the brief says only what its sources say. Each claim is matched back to the exact passage it draws on, and the facet weakens when statements drift beyond what the cited text supports. This is the facet that most directly guards against overstatement, so it carries substantial weight in the final index.
#Evidence strength
How strong is the underlying evidence? A causal finding from a controlled experiment supports a claim far more firmly than a correlation observed in a single dataset. This facet reflects the evidence type behind each claim — causal, associational, synthesis, or preclinical — and rewards briefs whose conclusions rest on the stronger end of that spectrum.
#Reproducibility
Has the finding been seen more than once? A result that independent groups have replicated is more dependable than a striking one-off. Reproducibility rises when a claim is corroborated across multiple independent sources and weakens when a conclusion hangs on a lone study.
#Durability
Does the finding still hold over time? Some results age well; others are overturned or quietly superseded as a field moves on. Durability considers whether a claim remains consistent with the more recent literature rather than resting on an isolated older report that later work has moved past.
#Reagent validity
Where it applies, were the key reagents or tools sound? Findings built on well-characterised, validated reagents — antibodies, cell lines, and the like — are more trustworthy than those resting on materials of unknown or questionable identity. For questions where reagents are not in play, this facet is marked N/A and does not affect the score.
#Trust bands
Alongside the number, each brief carries a trust band — a short, plain-language label such as High trust — so you can read the verdict at a glance without interpreting a raw score. The band tracks the index: higher scores map to higher-trust bands, and as facets slip toward Caution or Weak, the band steps down to signal that more of the brief deserves a human review before you act on it.
The band is a fast orientation, not a substitute for the detail. A mid-range band is an invitation to open the facets and see whether the soft spot sits somewhere that matters for how you intend to use the brief.
#How to read a low score
A low Trust Index is information, not a failure. It usually means the honest state of the literature on your question is thin, mixed, or fast-moving — and the agent is surfacing that rather than papering over it. The right response is to read why the score is low, not to discard the brief.
- Open the facets first. The mark pattern tells you the cause. A Weak reproducibility facet points to single-study claims; a Weak grounding facet is a stronger warning that some statements may reach past their sources.
- Trace the flagged claims. Every claim links to its source passage. Follow the citations behind the facets that scored low and read the underlying evidence yourself.
- Match the caution to your use. Weak durability may not matter for a background scan but matters a great deal before an experiment. Weigh the soft spot against the decision you are making.
- Keep the run going. Because a brief opens as a live research run, you can ask the agent to dig deeper on a specific low-scoring claim, pull in more sources, or narrow the question — and watch the facets and index update as the evidence firms up.
