Benchmarks

We build benchmarks for pharma-specific tasks that surface problems general models cannot solve — and we publish the method behind every number so you can re-run it.

Benchmarks

Internal comparisons, run August 2026 — same question, same time budget, scored the same way for every system. Model versions move monthly, so we rerun and republish when they do.

Factual claims supported by citations

How often a factual claim is backed by a source citation. Every citation in BIOPTIC also carries a verbatim supporting quotation from the source it came from.

BIOPTIC

75.2%

Gemini 3.1 Pro Deep Research

32.2%

GPT-5.5 (high)

31.3%

Claude Code with Opus 4.7 (high)

5.3%

Overall asset scouting quality

How well the system identifies qualifying assets while avoiding incorrect matches across complex, multi-criteria searches. Around half of the assets in this benchmark originate from Chinese and Japanese sources.

BIOPTIC

79.7%

Gemini 3.1 Deep Think

59.2%

Claude Opus 4.6 (high)

56.2%

OpenAI Deep Research

48.9%

Known competitors identified in the same indication

How often a known competing asset is surfaced when researching a therapeutic landscape.

BIOPTIC

83%

OpenAI Deep Research / o3-pro

65-67%

Perplexity Labs / Gemini 2.5 Pro

59-60%

How we measure

We take an evaluation-centric approach to building the system. The benchmarks above come out of it.

01

Scout

One question in, every matching asset out — including the ones no database groups together.

02

Evaluate

Red flags, comparables and the evidence behind each one, traced to the line.

03

Launch

Positioning, pathways and forecast factbooks, market by market.

04

Grow

Share, uptake and competitive response tracked as they move.

Each orbit is an evaluation set. Each card is one rubric criterion written by our specialists; the half-filled marker shows which side of the train/test split it sits on.

Publications

We are one of very few companies in this field that publishes. The benchmarks above come out of this work.

arXiv

BIOPTIC Agent Hunt Globally — Wide Search AI Agents for Drug Asset Scouting in Investing, Business Development, and Competitive Intelligence (arXiv, 2026)

Vlad Vinogradov, Alisa Vinogradova, Luba Greenwood, Ilya Yasny, Dmitry Kobyzev, Shoman Kasbekar, Kong Nguyen, Dmitrii Radkevich, Roman Doronin, Andrey Doronichev

New benchmark milestone. Our arXiv preprint shows how Bioptic Agent achieves 79.7% F1 on drug asset scouting—outperforming every major Deep Research baseline—by combining tree-based self-learning, multilingual coverage, and completeness-first retrieval.

AAAI Award

LLM-Based Agents for Competitive Landscape Mapping in Drug Asset Due Diligence

Vlad Vinogradov, Alisa Vinogradova, Dmitrii Radkevich, Ilya Yasny, Dmitry Kobyzev, Ivan Izmailov, Katsiaryna Yanchanka, Andrey Doronichev

We present a competitor-discovery agent for drug asset due diligence that, given an indication, identifies competing drugs and extracts their attributes — a task complicated by fragmented, paywalled, alias-heavy, and fast-changing data. Since no public benchmark exists, we built one by converting five years of biotech VC diligence memos into a structured evaluation corpus, and added an LLM-judge agent to filter false positives. Our system, Bioptic Agent, achieves 83% recall, beating OpenAI Deep Research (65%) and Perplexity Labs (60%). Deployed in production, it cut analyst turnaround time for competitive analysis from 2.5 days to ~3 hours (~20x).

Tell us about your project

Tell us about your research question and deadline. We’ll follow up to discuss the scope, deliverables, and timing.

Name

Company

Work email

Country

Phone number
Project details (optional)
Thanks for getting in touch.
We’ve received your inquiry and will follow up by email.
Your form wasn’t submitted. Please try again or email office@bioptic.io.