Key concepts
A handful of primitives underpin everything BioSkepsis does. Learn these five and the rest of the documentation follows: a research brief is what you get, an agent run is how it is produced, evidence types classify what is known, the Trust Index scores how much to rely on it, and the literature landscape maps the field around it.
#The research brief
The research brief is the primary output of BioSkepsis. You give the agent a single research question; it returns a fully worked, cited answer rather than a list of links. A brief is a structured document, not a chat reply, and it is the unit everything else attaches to.
A finished brief contains:
- A synthesized answer to your question, broken down across the sub-questions the agent planned.
- Findings — the individual, load-bearing claims that make up the answer, each labeled by evidence type and carrying its own confidence.
- Numbered citations linking every claim back to the exact passage in a source, marked as full-text-verified or abstract-only.
- A Trust Index from 0 to 100 with six facet checks.
- Explicit contradictions and evidence gaps, plus a literature landscape of the surrounding field.
Each sub-question in a brief is marked Covered, Limited evidence, or Not established, so you can see at a glance where the literature actually supports an answer and where it thins out.
#The agent run
The agent run is the process that produces a brief. BioSkepsis is agentic: it plans the work, executes each step, and shows its reasoning as it goes. There is no separate manual mode to toggle — the agent is the product, and every question triggers a run through the same transparent workflow.
- Plans the question. Decomposes your question into focused sub-questions and works through each in turn.
- Searches the literature. Searches broad life-science literature (40M+ papers) and weights authoritative, top-tier sources, matching on the biological concepts a paper discusses rather than keywords alone.
- Reads full text. Reads entire papers — methods, results, and discussion — not just abstracts, so it captures caveats and context.
- Weighs the evidence. Types each finding and weighs supporting against contradictory evidence, surfacing contradictions and gaps.
- Screens for safety. Blocks retracted papers and hijacked-journal sources and flags corrections before they reach your conclusions.
- Verifies and writes. Checks every claim against full text, then synthesizes the brief and computes its Trust Index.
A research notebook accompanies each run: it records the agent's reasoning, the sources it weighed, and the per-finding confidence behind the final answer. Because the run is transparent, you can open any stage and inspect the work rather than trusting a black box.
#Evidence types
Not all evidence is equal, so every finding is labeled by the kind of evidence behind it. The type tells you how far a claim can be pushed — a causal result supports a stronger conclusion than a correlation. The four types are:
| Type | What it means |
|---|---|
| Causal | An intervention or mechanism establishes cause and effect. |
| Associational | A correlation is reported, without established causation. |
| Synthesis | Drawn from reviews or meta-analyses across multiple studies. |
| Preclinical | From in vitro or animal models rather than clinical data. |
A finding backed by two or more independent sources is additionally tagged Synthesized, marking it as corroborated rather than resting on a single paper. Where the studies the agent reviewed disagree, the disagreement is surfaced as a contradiction rather than smoothed over; where the literature has not answered a question, it is surfaced as an evidence gap.
#The Trust Index
The Trust Index is a single 0-to-100 score, with a trust band such as High trust, that summarizes how much to rely on a brief. It is not a quality grade on the writing; it is a composite of how well the underlying evidence holds up. Briefs use this index and its facet checks — never letter grades.
The score is backed by six facet checks, each marked OK, Caution, Weak, or N/A:
- Provenance — are the sources reputable and traceable?
- Grounding — is each claim tied to a cited passage?
- Evidence strength — how strong is the underlying evidence?
- Reproducibility — has the finding been replicated?
- Durability — does it still hold over time?
- Reagent validity — are the key reagents or tools validated, where applicable?
Reading the facets matters as much as the number. A brief can score well overall while showing Caution on reproducibility, which tells you exactly where to apply your own judgment before acting on the answer.
#The literature landscape
The literature landscape is a map of the biology your question touches. Instead of treating papers as a flat list, BioSkepsis reduces each one to the biology it actually discusses — genes, proteins, drugs, diseases, pathways, and GO and MeSH terms — and connects papers whose biological fingerprints overlap. The result is a picture of the field, not just a set of results.
From that map the agent draws out structure:
- Themes — community detection groups connected papers into the research themes a field is organized around.
- Anchors and bridges — the foundational papers a theme rests on, and the papers that bridge one theme to another, are identified.
- Momentum — each theme is shown as rising or fading, so you can tell active areas from settled ones.
Connections are weighted by distinctive shared concepts, so closeness reflects genuine topical similarity rather than shared citation habits. The map is built from the same curated, full-text sources as the rest of the run and is woven, citation-verified, into your brief — so the landscape you see is grounded in the evidence, not a separate visualization bolted on afterward.
