We build benchmarks for pharma-specific tasks that surface problems general models cannot solve — and we publish the method behind every number so you can re-run it.
Benchmarks
Internal comparisons, run August 2026 — same question, same time budget, scored the same way for every system. Model versions move monthly, so we rerun and republish when they do.
Factual claims supported by citations
How often a factual claim is backed by a source citation. Every citation in BIOPTIC also carries a verbatim supporting quotation from the source it came from.
BIOPTIC
75.2%
Gemini 3.1 Pro Deep Research
32.2%
GPT-5.5 (high)
31.3%
Claude Code with Opus 4.7 (high)
5.3%
Overall asset scouting quality
How well the system identifies qualifying assets while avoiding incorrect matches across complex, multi-criteria searches. Around half of the assets in this benchmark originate from Chinese and Japanese sources.
How we measure
We take an evaluation-centric approach to building the system. The benchmarks above come out of it.
Scout
One question in, every matching asset out — including the ones no database groups together.
Evaluate
Red flags, comparables and the evidence behind each one, traced to the line.
Launch
Positioning, pathways and forecast factbooks, market by market.
Grow
Share, uptake and competitive response tracked as they move.

Each orbit is an evaluation set. Each card is one rubric criterion written by our specialists; the half-filled marker shows which side of the train/test split it sits on.
Publications
We are one of very few companies in this field that publishes. The benchmarks above come out of this work.
BIOPTIC Agent Hunt Globally — Wide Search AI Agents for Drug Asset Scouting in Investing, Business Development, and Competitive Intelligence (arXiv, 2026)
Vlad Vinogradov, Alisa Vinogradova, Luba Greenwood, Ilya Yasny, Dmitry Kobyzev, Shoman Kasbekar, Kong Nguyen, Dmitrii Radkevich, Roman Doronin, Andrey Doronichev
New benchmark milestone. Our arXiv preprint shows how Bioptic Agent achieves 79.7% F1 on drug asset scouting—outperforming every major Deep Research baseline—by combining tree-based self-learning, multilingual coverage, and completeness-first retrieval.
LLM-Based Agents for Competitive Landscape Mapping in Drug Asset Due Diligence
Vlad Vinogradov, Alisa Vinogradova, Dmitrii Radkevich, Ilya Yasny, Dmitry Kobyzev, Ivan Izmailov, Katsiaryna Yanchanka, Andrey Doronichev
We present a competitor-discovery agent for drug asset due diligence that, given an indication, identifies competing drugs and extracts their attributes — a task complicated by fragmented, paywalled, alias-heavy, and fast-changing data. Since no public benchmark exists, we built one by converting five years of biotech VC diligence memos into a structured evaluation corpus, and added an LLM-judge agent to filter false positives. Our system, Bioptic Agent, achieves 83% recall, beating OpenAI Deep Research (65%) and Perplexity Labs (60%). Deployed in production, it cut analyst turnaround time for competitive analysis from 2.5 days to ~3 hours (~20x).