How the agent works
Give the agent one research question and it runs the whole pipeline end to end — planning, searching, reading full text, weighing evidence, screening for reliability, and returning a cited brief. There is no mode to switch on: this hands-off run is simply how BioSkepsis works.
#The end-to-end run
You start a run the same way every time: type a research question and let the agent take it from there. What follows is a single, continuous pipeline rather than a set of tools you operate step by step. The agent plans the question, chooses and issues searches, filters results for relevance, selects the strongest papers, reads them in full, weighs what they actually show, screens every source for reliability, and synthesizes a brief in which each claim is tied back to the passage that supports it.
Because the run is autonomous, you are not asked to pick queries, skim result lists, or decide which PDFs are worth opening. The agent makes those choices the way a careful researcher would — and, just as importantly, it shows its work so you can follow and correct it. The sections below walk through each stage in the order the agent performs them.
#Plans the question
The agent begins by breaking your question into focused sub-questions and deciding how to approach each one. A broad prompt like “does metformin have anticancer effects?” becomes a small set of tractable lines of inquiry — mechanism, preclinical evidence, clinical outcomes, and known contradictions — that can each be searched and answered on its own.
This planning step is what lets a single question drive a thorough run. Instead of firing one keyword search and hoping for the best, the agent maps out what it needs to establish before it can answer you, then works through each sub-question in turn and carries the results forward into the final synthesis.
#Curated search
For each sub-question the agent composes and issues its own searches across broad life-science literature — 40M+ papers spanning the biomedical and life sciences. It does not rely on you to supply the right keywords. It reasons about the biological concepts a question actually touches, so relevant work is not missed just because a paper uses different terminology than your prompt.
Retrieved results are then filtered for relevance and weighted toward authoritative, top-tier sources rather than treated as an undifferentiated pile of links. From that pool the agent selects the strongest papers to read in full — the equivalent of a researcher triaging a search return down to the handful of studies actually worth their time.
#Full-text reading
This is the stage that separates a real research run from an abstract summary. Rather than stopping at the abstract, the agent reads each selected paper in full — introduction, methods, results, and discussion. Abstracts routinely omit sample sizes, model systems, dosing, and the caveats authors reserve for the discussion; reading the full text is how the agent captures the context and limitations that determine what a finding really means.
Full-text reading also grounds everything downstream. Because the agent has the actual passages in hand, it can weigh evidence on what a study did rather than on what its title implies, and it can later verify each claim against the exact sentence that supports it.
#Evidence weighing
With the papers read, the agent evaluates what they collectively show. It types each finding — distinguishing causal claims from associations, from broader syntheses, from preclinical results — and weighs supporting evidence against contradictory evidence rather than counting papers that agree.
Disagreement is surfaced, not smoothed over. When studies conflict, the agent flags the contradiction and characterizes both sides; where the literature is thin or silent, it names the evidence gap instead of papering over it. The result is a picture of not just what the field claims, but how well those claims are actually supported.
#Trust & safety
Before any source is allowed to shape your conclusions, the agent screens it. Retracted papers and known hijacked-journal sources are blocked outright, so a discredited study cannot quietly prop up a claim. Papers carrying corrections or expressions of concern are flagged rather than silently dropped, so you can see when a caveat applies.
This screening runs as part of every pipeline, not as an optional check. It is the reason the agent can hand you a brief you can act on: the evidence behind it has already been filtered for the failure modes that most often mislead an automated literature review.
#The cited brief
The run ends by synthesizing everything into a single readable brief that answers your original question. Before a sentence makes it into that brief, the agent verifies the claim against the full text of its source — every claim links back to the exact passage it rests on, so nothing is asserted without a traceable basis.
Alongside the answer, the brief carries a Trust Index that scores how well the conclusion is supported, and it keeps the contradictions and evidence gaps the agent found in view rather than burying them. You get a grounded, verifiable answer instead of a list of links — and you can trace any part of it back to the source in one click.
#Watching and steering a run
A run is not a black box. Every stage is transparent as it happens: you can watch the agent plan the question, issue searches, select and read papers, and weigh the evidence, and you can open the reasoning behind any step to see why a paper was chosen or a finding was typed the way it was.
When the run finishes, the brief is a starting point you can push on. Review the sources the agent gathered, add or remove papers, and ask targeted follow-up questions to go deeper on a mechanism, a contradiction, or a gap. Each follow-up continues the same run, so the agent builds on the evidence it has already read rather than starting over.
