Adding patents to Paperclip
Most publicly disclosed drug compounds are first reported in patents, not papers. Patent disclosures also frequently precede corresponding reports in the scientific literature by several years. However, patents are written to establish legal claims, which makes their scientific content difficult to analyze systematically. Paperclip now provides access to a global catalog of 170 million patent publications from 75 jurisdictions. Within this catalog, 1.8 million drug-related patents include full-text claims and descriptions, together with 30 million compound annotations. Agents can search and read this information alongside scientific papers, regulatory documents, and clinical trials.
What's included
- Drug-discovery subset: 10,377,132 publications selected from the global catalog by 468 curated CPC classification codes covering small-molecule, biologic, and diagnostic patent families.
- Full text: 1,842,503 USPTO and CNIPA publications with claims and descriptions. USPTO coverage through April 2026. CNIPA full text since October 2023.
- Compound annotations: 30,990,818 SureChEMBL compounds linked to individual patents and queryable via SQL.
- Sequences: 294,678,971 sequences from the USPTO Patent Sequence Information Products Suite (PSIPS) overflow files across 21,585 patent documents, plus 925,941 CNIPA ST.26 sequences. PSIPS covers sequences submitted as separate sequence listings; additional sequences may appear inline in patent descriptions.
- Global catalog: 170.4 million patent records with localized titles, abstracts, and metadata (classification codes, parties, dates, families).
Already have Paperclip? Patents are included by default. Run paperclip update and ask your agent to search patents. For example: “Find a patent on GLP-1 receptor agonists.”
New to Paperclip? Install with curl -fsSL https://paperclip.gxl.ai/install.sh | bash, or add it as an MCP server at https://paperclip.gxl.ai/mcp.
See the full documentation for details.
Why This Matters
Patent filings contain data like compound structures, structure-activity relationship (SAR) tables, and antibody sequences, disclosed because patent law requires it. That data can be the first public record of a program, or useful experimental detail that would not have been published otherwise.
However, patents are written in language designed for legal breadth rather than scientific clarity. Information relevant to drug discovery is often buried in description sections. Paperclip indexes this into the same virtual filesystem it uses for the rest of its biomedical corpus. An agent can grep for a specific SMILES string or target name across 1.8 million patent descriptions and get back a line number, rather than fetching and searching through each document individually.
Below, we show three case studies: a target landscape analysis across PCSK9 and SARM1, an antibody-oligonucleotide conjugation (AOC) chemistry landscape, and an antibody developability deep-dive on anti-IL-13 antibodies from patent data.
Case Study: Target Landscape Analysis
We ran a target landscape analysis on two targets (PCSK9 and SARM1), comparing an agent with Paperclip against an agent with web search. Both used Claude Code with Opus 5.
Find small-molecule chemical matter targeting {TARGET}. Compounds in patents and any tool compounds in published literature. Cover all modality directions. For each compound or series: name, owner, patent or paper ID, modality, development stage, and potency on record.
| Opus 5 + Paperclip | Opus 5 + Web search | |
|---|---|---|
| PCSK9 compound rows | 35 | 25 |
| SARM1 compound rows | 39 | 19 |
The agent with Paperclip found more compounds across both targets, and more of them came with patent citations and potency data. The difference was especially clear in what each agent could pull from inside patent descriptions.
For the PCSK9 analysis, the agent with Paperclip grepped an AstraZeneca patent (US-20240228469-A1) and pulled the activity profile for a distinct 2022-priority AstraZeneca patent family. The agent with web search found AstraZeneca’s compound AZD0780 from its 2019-priority family, which is publicly known, but not this series.
| Assay | Value |
|---|---|
| TR-FRET probe displacement IC50 | 0.7 nM |
| BIAcore Kd | 0.3 nM |
| LDL-uptake potency | 0.1 μM |
| hERG IC50 (off-target) | 12.0 μM |
| GSK3β IC50 (off-target) | >10 μM |
The SARM1 analysis looked similar. The agent with Paperclip pulled activity data straight from patent descriptions that the agent with web search did not, such as a Sironax table covering compounds 1–120 (US-20250236613-A1) and a Chinese patent assigned to Kehui Zhiyao with SARM1 enzyme assay data (CN-119219603-A).
| Compound | Owner | Patent / paper | Potency |
|---|---|---|---|
| DSRM-3716 (isoquinoline) | Disarm / Wash. Univ. | WO-2018057989-A1; PMC8179325 | IC50 75 nM (biochemical); 2.8 µM (cell cADPR) |
| Sironax compounds 1–120 | Sironax Ltd | US-20250236613-A1 | IC50 bins: A <5 µM; B 5–<15 µM; C 15–30 µM; D >30 µM |
| Nico Therapeutics (2022 family) | Nico Therapeutics | US-20250361234-A1 | IC50 bins: +++ <1 µM; ++ 1–10 µM; + >10 µM |
| Kehui Zhiyao SARM1 inhibitor series | Kehui Zhiyao (Shenzhen) New Drug Research Center | CN-119219603-A | IC50 bins: A <1 µM; B 1–10 µM; C 10–30 µM |
| Compound 331P1 | Not stated | PMC11428815 | IC50 189.3 nM |
| Compound | Owner | Patent / paper | Potency |
|---|---|---|---|
| DSRM-3716 (isoquinoline) | Disarm / Eli Lilly | PMC8179325 | IC50 75 nM |
| Disarm bicyclic/tricyclic genus | Disarm / Eli Lilly | WO2019236879A1 | "No quantitative IC50 values provided" |
| Disarm 2nd-gen heteroaryl | Disarm / Eli Lilly | WO2021142006A1 | "Does not contain numerical potency values" |
| LY3873862 | Eli Lilly | — (review article only) | UNKNOWN |
| NB-4746 | Nura Bio | — (press release only) | UNKNOWN |
The agent with Paperclip also found that AstraZeneca had discontinued further development of its BEXi compound series. Sub-inhibitory concentrations activated SARM1 and accelerated neurodegeneration: “These results prompted us to discontinue BEXi development in favour of alternative strategies.” It surfaced from a November 2025 bioRxiv preprint during a patent landscape analysis, because patents, papers, and trials are all in the same virtual filesystem. The agent with web search did not find it.
Chinese Patent Coverage
We also ran a China-focused analysis asking for Chinese (CN) patents covering small-molecule chemical matter targeting PCSK9 and SARM1.
| Opus 5 + Paperclip | Opus 5 + Web search | |
|---|---|---|
| PCSK9 patents found | 21 | 7 |
| SARM1 patents found | 34 | 8 |
Paperclip indexes 250K China National Intellectual Property Administration (CNIPA) full-text descriptions with line-numbered claims, while web search has no structured access to Chinese patent text.
The difference showed up clearly when mapping Chinese applications to publication numbers. The agent with web search found the Salubris Pharmaceuticals family via the US grant and the Chinese priority application, then spent seven tool calls trying to map the application number to a CN publication, searching 信立泰 + PCSK9, guessing a 2018 publication number that turned out to be a different filing, and eventually giving up. The agent with Paperclip searched the CN patent index and returned CN110546149A and its grant, CN110546149B, directly.
The analysis also exposed a limitation: 17 of 34 SARM1 CN entries had an unknown chemotype, likely because compound data in those patents sits in images or Markush structures that SureChEMBL does not extract.
Case Study: Antibody-Oligonucleotide Conjugation Landscape
The first case study focused on finding compounds for a specific target. The next question was broader: we ran the same comparison on the antibody-oligonucleotide conjugation (AOC) chemistry landscape.
I want to understand chemical conjugation of oligonucleotides to antibodies for therapeutics. Tabulate the key strengths/weaknesses of each approach from a pharmaceutical developability perspective, covering topics like chemical stability, in vivo stability/cleavability, manufacturability, etc. Then, indicate who’s innovating in each strategy based on the patent (or peer-reviewed) literature. Summarize what they’ve published on.
- Found Avidity, Dyne, Denali, Genentech/Roche, Tallac, Alnylam, Regeneron, Sorrento, Sapreme Technologies, Gennao Bio, GeneQuantum, Janssen, Code Biotherapeutics, and others
- Named 25 patent identifiers inline
- Found Avidity, Dyne, Denali, Genentech/Roche, Tallac, Takeda, and CrossBridge
- Named 1 patent identifier inline
According to a domain expert, both covered the conjugation chemistry and linker trade-offs comparably, but Paperclip was significantly more comprehensive in identifying innovators and citing the patents behind them.
Alnylam, one of the largest companies in oligonucleotide therapeutics, did not appear in the web search results.
The agent with Paperclip found Sapreme Technologies through a patent titled “Antibody-oligonucleotide conjugate” (US-12453782-B2), which is about as generic as a patent title gets. A grep of the full text revealed what they’re actually working on:
Saponins on the conjugate to help the payload escape the endosome, which is a major delivery bottleneck for AOCs. The abstract mentions saponin, but the mechanism and the experimental evidence supporting it are only in the patent description. The agent with web search did not surface Sapreme.
Interestingly, when prompted to reflect on its own output, the agent with web search acknowledged the gap: “patent-only actors, big pharma internal programs, Chinese biotechs, tool/linker companies, are systematically underrepresented.”
Case Study: Antibody Developability from Patent Data
The first two case studies focused on competitive landscape. Patent descriptions also disclose the specific problems that arose during development, and what solved them.
We asked Paperclip to find developability data for anti-IL-13 antibodies across patent filings, and it found quantitative data in 13 of 24 relevant patents, covering tralokinumab (AstraZeneca), SAR156597 (Sanofi, an anti-IL-4/IL-13 bispecific), lebrikizumab (Genentech), and Construct 133 (Apogee Therapeutics). Here are some of the more interesting rows it pulled:
| Antibody | Property | Value | From the patent |
|---|---|---|---|
| SAR156597 | Aggregation | 0.5-1% HMW/hr at 25°C | "this antibody has such a strong propensity to aggregate that it cannot be formulated in a liquid in the concentration range targeted." Low ionic strength at pH 7.0 with proline slowed aggregation. (Patent ID: US-10005835-B2) |
| SAR156597 | Solubility | pI 5.8-6.2 | "the anti-IL4/anti-IL13 bispecific antibody has a particularly low isoelectric point, making it more difficult to formulate due to solubility issues." pH 6.2 caused precipitation (histidine) or gel formation (succinate). (Patent ID: US-10005835-B2) |
| Tralokinumab | Viscosity | Target specification: <10 cP at 23°C | "One of the major challenges...is the high viscosity of anti-IL13 antibodies, such as Tralokinumab, at high concentrations, e.g. 150 mg/mL." 150 mM lysine "dramatically decreased the viscosity." (Patent ID: US-20240336678-A1) |
| IgG1 variants (including Construct 133) | Aggregation | <15% after Protein A | "IgG4 variants had an opalescent appearance...and an aggregation sensitivity (>68% aggregate)." IgG1 variants "remained a clear solution with <15% aggregate." (Patent ID: US-12358979-B2) |
| Construct 133 | Half-life | 27.6 days (NHP) | YTE Fc substitution. Lebrikizumab = 18 days NHP. Predicted 80-110 days in humans. (Patent ID: US-12358979-B2) |
Apogee’s US-12358979-B2 reports thermal stability and aggregation data across its engineered anti-IL-13 panel, and compares Construct 133 head-to-head with lebrikizumab on half-life and clearance. Construct 133 showed a 27.6-day NHP half-life versus 18 days for lebrikizumab, with a predicted 80-110-day human half-life. That kind of engineering narrative lives in the patent description.
Conclusion
Patent data is one of the richest sources in drug discovery, and historically one of the hardest to search programmatically. The case studies above show what changes when an agent can search and grep across 1.8 million patent descriptions in the same virtual filesystem as the rest of the biomedical literature: compound series that web search did not find, a program discontinuation surfaced from a preprint during a patent landscape analysis, and developability data extracted verbatim from patent descriptions. Paperclip makes all of it agent-native alongside the rest of its biomedical corpus.