gxl
Blog
August 25, 2026

Adding patents to Paperclip

TL;DR

Most publicly disclosed drug compounds are first reported in patents, not papers. Patent disclosures also frequently precede corresponding reports in the scientific literature by several years. However, patents are written to establish legal claims, which makes their scientific content difficult to analyze systematically. Paperclip now provides access to a global catalog of 170 million patent publications from 75 jurisdictions. Within this catalog, 1.8 million drug-related patents include full-text claims and descriptions, together with 30 million compound annotations. Agents can search and read this information alongside scientific papers, regulatory documents, and clinical trials.

170M+
Patent Titles & Abstracts
1.8M+
Full-Text Claims & Descriptions
30M+
Compound Annotations
75
Jurisdictions: US, China, Europe & More
What's included
  • Drug-discovery subset: 10,377,132 publications selected from the global catalog by 468 curated CPC classification codes covering small-molecule, biologic, and diagnostic patent families.
  • Full text: 1,842,503 USPTO and CNIPA publications with claims and descriptions. USPTO coverage through April 2026. CNIPA full text since October 2023.
  • Compound annotations: 30,990,818 SureChEMBL compounds linked to individual patents and queryable via SQL.
  • Sequences: 294,678,971 sequences from the USPTO Patent Sequence Information Products Suite (PSIPS) overflow files across 21,585 patent documents, plus 925,941 CNIPA ST.26 sequences. PSIPS covers sequences submitted as separate sequence listings; additional sequences may appear inline in patent descriptions.
  • Global catalog: 170.4 million patent records with localized titles, abstracts, and metadata (classification codes, parties, dates, families).
How to use it

Already have Paperclip? Patents are included by default. Run paperclip update and ask your agent to search patents. For example: “Find a patent on GLP-1 receptor agonists.”

New to Paperclip? Install with curl -fsSL https://paperclip.gxl.ai/install.sh | bash, or add it as an MCP server at https://paperclip.gxl.ai/mcp.

See the full documentation for details.

Why This Matters

Patent filings contain data like compound structures, structure-activity relationship (SAR) tables, and antibody sequences, disclosed because patent law requires it. That data can be the first public record of a program, or useful experimental detail that would not have been published otherwise.

However, patents are written in language designed for legal breadth rather than scientific clarity. Information relevant to drug discovery is often buried in description sections. Paperclip indexes this into the same virtual filesystem it uses for the rest of its biomedical corpus. An agent can grep for a specific SMILES string or target name across 1.8 million patent descriptions and get back a line number, rather than fetching and searching through each document individually.

$ ls /patents/US-11633476-B2/
meta.json Metadata, dates, assignees, family
content.lines Full text (claims + description)
surechembl/ Linked compounds (SMILES, InChIKey, MW)
sequences.tsv Patent sequence listings
$ paperclip search -s patents "SARM1 small molecule inhibitor"
Found 12 patents [s_a1b2c3d4]
1. Small molecule modulators of SARM1
US-2021261537-A1 | Disarm Therapeutics
2. SARM1 inhibitors
CN-119219603-A | Kehui Zhiyao (Shenzhen)
...
$ paperclip grep "IC50" /patents/CN-119219603-A/content.lines
L522: 生物活性实施例2:抑制SARM1酶活性的体外生物化学测试(IC50):
L527: 下表1中提供了这些化合物在测定中的IC50区间:
L528: 抑制SARM1酶活性的IC50区间:A<1μM;B:1-10μM;C:10-30μM

Below, we show three case studies: a target landscape analysis across PCSK9 and SARM1, an antibody-oligonucleotide conjugation (AOC) chemistry landscape, and an antibody developability deep-dive on anti-IL-13 antibodies from patent data.

Case Study: Target Landscape Analysis

We ran a target landscape analysis on two targets (PCSK9 and SARM1), comparing an agent with Paperclip against an agent with web search. Both used Claude Code with Opus 5.

QUERY

Find small-molecule chemical matter targeting {TARGET}. Compounds in patents and any tool compounds in published literature. Cover all modality directions. For each compound or series: name, owner, patent or paper ID, modality, development stage, and potency on record.

RESULTS
Compound landscape coverage for PCSK9 and SARM1.
Opus 5 + PaperclipOpus 5 + Web search
PCSK9 compound rows3525
SARM1 compound rows3919

The agent with Paperclip found more compounds across both targets, and more of them came with patent citations and potency data. The difference was especially clear in what each agent could pull from inside patent descriptions.

For the PCSK9 analysis, the agent with Paperclip grepped an AstraZeneca patent (US-20240228469-A1) and pulled the activity profile for a distinct 2022-priority AstraZeneca patent family. The agent with web search found AstraZeneca’s compound AZD0780 from its 2019-priority family, which is publicly known, but not this series.

From US-20240228469-A1, Example 116, Tables 11–12 (Opus 5 + Paperclip)
AssayValue
TR-FRET probe displacement IC500.7 nM
BIAcore Kd0.3 nM
LDL-uptake potency0.1 μM
hERG IC50 (off-target)12.0 μM
GSK3β IC50 (off-target)>10 μM

The SARM1 analysis looked similar. The agent with Paperclip pulled activity data straight from patent descriptions that the agent with web search did not, such as a Sironax table covering compounds 1–120 (US-20250236613-A1) and a Chinese patent assigned to Kehui Zhiyao with SARM1 enzyme assay data (CN-119219603-A).

RESULTS
SARM1 compound coverage. The Paperclip agent surfaced patent-specific series from Sironax and Kehui Zhiyao not found by web search.
Opus 5 + Paperclip (5 of 39 rows)
CompoundOwnerPatent / paperPotency
DSRM-3716 (isoquinoline)Disarm / Wash. Univ.WO-2018057989-A1; PMC8179325IC50 75 nM (biochemical); 2.8 µM (cell cADPR)
Sironax compounds 1–120Sironax LtdUS-20250236613-A1IC50 bins: A <5 µM; B 5–<15 µM; C 15–30 µM; D >30 µM
Nico Therapeutics (2022 family)Nico TherapeuticsUS-20250361234-A1IC50 bins: +++ <1 µM; ++ 1–10 µM; + >10 µM
Kehui Zhiyao SARM1 inhibitor seriesKehui Zhiyao (Shenzhen) New Drug Research CenterCN-119219603-AIC50 bins: A <1 µM; B 1–10 µM; C 10–30 µM
Compound 331P1Not statedPMC11428815IC50 189.3 nM
Opus 5 + Web search (5 of 19 rows)
CompoundOwnerPatent / paperPotency
DSRM-3716 (isoquinoline)Disarm / Eli LillyPMC8179325IC50 75 nM
Disarm bicyclic/tricyclic genusDisarm / Eli LillyWO2019236879A1"No quantitative IC50 values provided"
Disarm 2nd-gen heteroarylDisarm / Eli LillyWO2021142006A1"Does not contain numerical potency values"
LY3873862Eli Lilly— (review article only)UNKNOWN
NB-4746Nura Bio— (press release only)UNKNOWN

The agent with Paperclip also found that AstraZeneca had discontinued further development of its BEXi compound series. Sub-inhibitory concentrations activated SARM1 and accelerated neurodegeneration: “These results prompted us to discontinue BEXi development in favour of alternative strategies.” It surfaced from a November 2025 bioRxiv preprint during a patent landscape analysis, because patents, papers, and trials are all in the same virtual filesystem. The agent with web search did not find it.

Chinese Patent Coverage

We also ran a China-focused analysis asking for Chinese (CN) patents covering small-molecule chemical matter targeting PCSK9 and SARM1.

RESULTS
Chinese patent landscape coverage for PCSK9 and SARM1.
Opus 5 + PaperclipOpus 5 + Web search
PCSK9 patents found217
SARM1 patents found348

Paperclip indexes 250K China National Intellectual Property Administration (CNIPA) full-text descriptions with line-numbered claims, while web search has no structured access to Chinese patent text.

The difference showed up clearly when mapping Chinese applications to publication numbers. The agent with web search found the Salubris Pharmaceuticals family via the US grant and the Chinese priority application, then spent seven tool calls trying to map the application number to a CN publication, searching 信立泰 + PCSK9, guessing a 2018 publication number that turned out to be a different filing, and eventually giving up. The agent with Paperclip searched the CN patent index and returned CN110546149A and its grant, CN110546149B, directly.

The analysis also exposed a limitation: 17 of 34 SARM1 CN entries had an unknown chemotype, likely because compound data in those patents sits in images or Markush structures that SureChEMBL does not extract.

Case Study: Antibody-Oligonucleotide Conjugation Landscape

The first case study focused on finding compounds for a specific target. The next question was broader: we ran the same comparison on the antibody-oligonucleotide conjugation (AOC) chemistry landscape.

QUERY

I want to understand chemical conjugation of oligonucleotides to antibodies for therapeutics. Tabulate the key strengths/weaknesses of each approach from a pharmaceutical developability perspective, covering topics like chemical stability, in vivo stability/cleavability, manufacturability, etc. Then, indicate who’s innovating in each strategy based on the patent (or peer-reviewed) literature. Summarize what they’ve published on.

RESULTS
Opus 5 + Paperclip
  • Found Avidity, Dyne, Denali, Genentech/Roche, Tallac, Alnylam, Regeneron, Sorrento, Sapreme Technologies, Gennao Bio, GeneQuantum, Janssen, Code Biotherapeutics, and others
  • Named 25 patent identifiers inline
Opus 5 + Web search
  • Found Avidity, Dyne, Denali, Genentech/Roche, Tallac, Takeda, and CrossBridge
  • Named 1 patent identifier inline

According to a domain expert, both covered the conjugation chemistry and linker trade-offs comparably, but Paperclip was significantly more comprehensive in identifying innovators and citing the patents behind them.

Alnylam, one of the largest companies in oligonucleotide therapeutics, did not appear in the web search results.

The agent with Paperclip found Sapreme Technologies through a patent titled “Antibody-oligonucleotide conjugate” (US-12453782-B2), which is about as generic as a patent title gets. A grep of the full text revealed what they’re actually working on:

$ paperclip grep -m 8 "endosomal escape|saponin|cleavable|conjugat" /patents/US-12453782-B2/content.lines
L2: The invention relates to a ligand-effector moiety provided with at
least one saponin and to antibody-effector moiety provided with at
least one saponin, such as an antibody-drug conjugate and an
antibody-oligonucleotide conjugate. ... The invention also relates
to an antibody-drug conjugate comprising covalently linked saponin
and to an antibody-oligonucleotide conjugate comprising covalently
linked saponin.

Saponins on the conjugate to help the payload escape the endosome, which is a major delivery bottleneck for AOCs. The abstract mentions saponin, but the mechanism and the experimental evidence supporting it are only in the patent description. The agent with web search did not surface Sapreme.

Interestingly, when prompted to reflect on its own output, the agent with web search acknowledged the gap: “patent-only actors, big pharma internal programs, Chinese biotechs, tool/linker companies, are systematically underrepresented.”

Case Study: Antibody Developability from Patent Data

The first two case studies focused on competitive landscape. Patent descriptions also disclose the specific problems that arose during development, and what solved them.

We asked Paperclip to find developability data for anti-IL-13 antibodies across patent filings, and it found quantitative data in 13 of 24 relevant patents, covering tralokinumab (AstraZeneca), SAR156597 (Sanofi, an anti-IL-4/IL-13 bispecific), lebrikizumab (Genentech), and Construct 133 (Apogee Therapeutics). Here are some of the more interesting rows it pulled:

RESULTS
Developability data extracted from anti-IL-13 antibody patent filings.
AntibodyPropertyValueFrom the patent
SAR156597Aggregation0.5-1% HMW/hr at 25°C"this antibody has such a strong propensity to aggregate that it cannot be formulated in a liquid in the concentration range targeted." Low ionic strength at pH 7.0 with proline slowed aggregation. (Patent ID: US-10005835-B2)
SAR156597SolubilitypI 5.8-6.2"the anti-IL4/anti-IL13 bispecific antibody has a particularly low isoelectric point, making it more difficult to formulate due to solubility issues." pH 6.2 caused precipitation (histidine) or gel formation (succinate). (Patent ID: US-10005835-B2)
TralokinumabViscosityTarget specification: <10 cP at 23°C"One of the major challenges...is the high viscosity of anti-IL13 antibodies, such as Tralokinumab, at high concentrations, e.g. 150 mg/mL." 150 mM lysine "dramatically decreased the viscosity." (Patent ID: US-20240336678-A1)
IgG1 variants (including Construct 133)Aggregation<15% after Protein A"IgG4 variants had an opalescent appearance...and an aggregation sensitivity (>68% aggregate)." IgG1 variants "remained a clear solution with <15% aggregate." (Patent ID: US-12358979-B2)
Construct 133Half-life27.6 days (NHP)YTE Fc substitution. Lebrikizumab = 18 days NHP. Predicted 80-110 days in humans. (Patent ID: US-12358979-B2)

Apogee’s US-12358979-B2 reports thermal stability and aggregation data across its engineered anti-IL-13 panel, and compares Construct 133 head-to-head with lebrikizumab on half-life and clearance. Construct 133 showed a 27.6-day NHP half-life versus 18 days for lebrikizumab, with a predicted 80-110-day human half-life. That kind of engineering narrative lives in the patent description.

Conclusion

Patent data is one of the richest sources in drug discovery, and historically one of the hardest to search programmatically. The case studies above show what changes when an agent can search and grep across 1.8 million patent descriptions in the same virtual filesystem as the rest of the biomedical literature: compound series that web search did not find, a program discontinuation surfaced from a preprint during a patent landscape analysis, and developability data extracted verbatim from patent descriptions. Paperclip makes all of it agent-native alongside the rest of its biomedical corpus.