
Resources
Guides, research, and perspectives on R&D intelligence, IP strategy, and the future of AI enabled innovation.

Executive Summary
In 2024, US patent infringement jury verdicts totaled $4.19 billion across 72 cases. Twelve individual verdicts exceeded $100million. The largest single award—$857 million in General Access Solutions v.Cellco Partnership (Verizon)—exceeded the annual R&D budget of many mid-market technology companies. In the first half of 2025 alone, total damages reached an additional $1.91 billion.
The consequences of incomplete patent intelligence are not abstract. In what has become one of the most instructive IP disputes in recent history, Masimo’s pulse oximetry patents triggered a US import ban on certain Apple Watch models, forcing Apple to disable its blood oxygen feature across an entire product line, halt domestic sales of affected models, invest in a hardware redesign, and ultimately face a $634 million jury verdict in November 2025. Apple—a company with one of the most sophisticated intellectual property organizations on earth—spent years in litigation over technology it might have designed around during development.
For organizations with fewer resources than Apple, the risk calculus is starker. A mid-size materials company, a university spinout, or a defense contractor developing next-generation battery technology cannot absorb a nine-figure verdict or a multi-year injunction. For these organizations, the patent landscape analysis conducted during the development phase is the primary risk mitigation mechanism. The quality of that analysis is not a matter of convenience. It is a matter of survival.
And yet, a growing number of R&D and IP teams are conducting that analysis using general-purpose AI tools—ChatGPT, Claude, Microsoft Co-Pilot—that were never designed for patent intelligence and are structurally incapable of delivering it.
This report presents the findings of a controlled comparison study in which identical patent landscape queries were submitted to four AI-powered tools: Cypris (a purpose-built R&D intelligence platform),ChatGPT (OpenAI), Claude (Anthropic), and Microsoft Co-Pilot. Two technology domains were tested: solid-state lithium-sulfur battery electrolytes using garnet-type LLZO ceramic materials (freedom-to-operate analysis), and bio-based polyamide synthesis from castor oil derivatives (competitive intelligence).
The results reveal a significant and structurally persistent gap. In Test 1, Cypris identified over 40 active US patents and published applications with granular FTO risk assessments. Claude identified 12. ChatGPT identified 7, several with fabricated attribution. Co-Pilot identified 4. Among the patents surfaced exclusively by Cypris were filings rated as “Very High” FTO risk that directly claim the technology architecture described in the query. In Test 2, Cypris cited over 100 individual patent filings with full attribution to substantiate its competitive landscape rankings. No general-purpose model cited a single patent number.
The most active sectors for patent enforcement—semiconductors, AI, biopharma, and advanced materials—are the same sectors where R&D teams are most likely to adopt AI tools for intelligence workflows. The findings of this report have direct implications for any organization using general-purpose AI to inform patent strategy, competitive intelligence, or R&D investment decisions.

1. Methodology
A controlled comparative evaluation was conducted on March 27, 2026. An identical patent landscape query was submitted verbatim to each platform under standardized testing conditions. No follow-up prompts, clarifications, or iterative refinements were permitted, ensuring that each platform was evaluated based solely on its initial response.
The outputs were preserved in their original form and evaluated against predefined criteria using publicly verifiable patent records.
1.1 Query
Identify all active US patents and published applications filed in the last 5 years related to solid-state lithium-sulfur battery electrolytes using garnet-type ceramic materials. For each, provide the assignee, filing date, key claims, and current legal status. Highlight any patents that could pose freedom-to-operate risks for a company developing a Li₇La₃Zr₂O₁₂(LLZO)-based composite electrolyte with a polymer interlayer.
1.2 Tools Evaluated

1.3 Evaluation Criteria
Each response was evaluated using a consistent six-part scoring framework: patent coverage, assignee accuracy, filing metadata completeness, depth of claim analysis, quality of FTO risk stratification, and the presence of actionable strategic guidance.
Patent numbers, assignees, filing information, and legal status were independently checked against publicly available USPTO and WIPO records. The evaluation focused on the completeness, accuracy, and practical utility of each platform’s output rather than writing quality or presentation.
2. Findings
2.1 Coverage Gap
The most significant finding is the scale of the coverage differential. Cypris identified over 40 active US patents and published applications spanning LLZO-polymer composite electrolytes, garnet interface modification, polymer interlayer architectures, lithium-sulfur specific filings, and adjacent ceramic composite patents. The results were organized by technology category with per-patent FTO risk ratings.
Claude identified 12 patents organized in a four-tier risk framework. Its analysis was structurally sound and correctly flagged the two highest-risk filings (Solid Energies US 11,967,678 and the LLZO nanofiber multilayer US 11,923,501). It also identified the University ofMaryland/ Wachsman portfolio as a concentration risk and noted the NASA SABERS portfolio as a licensing opportunity. However, it missed the majority of the landscape, including the entire Corning portfolio, GM's interlayer patents, theKorea Institute of Energy Research three-layer architecture, and the HonHai/SolidEdge lithium-sulfur specific filing.
ChatGPT identified 7 patents, but the quality of attribution was inconsistent. It listed assignees as "Likely DOE /national lab ecosystem" and "Likely startup / defense contractor cluster" for two filings—language that indicates the model was inferring rather than retrieving assignee data. In a freedom-to-operate context, an unverified assignee attribution is functionally equivalent to no attribution, as it cannot support a licensing inquiry or risk assessment.
Co-Pilot identified 4 US patents. Its output was the most limited in scope, missing the Solid Energies portfolio entirely, theUMD/ Wachsman portfolio, Gelion/ Johnson Matthey, NASA SABERS, and all Li-S specific LLZO filings.
2.2 Critical Patents Missed by Public Models
The following table presents patents identified exclusively by Cypris that were rated as High or Very High FTO risk for the proposed technology architecture. None were surfaced by any general-purpose model.

2.3 Patent Fencing: The Solid Energies Portfolio
Cypris identified a coordinated patent fencing strategy by Solid Energies, Inc. that no general-purpose model detected at scale. Solid Energies holds at least four granted US patents and one published application covering LLZO-polymer composite electrolytes across compositions(US-12463245-B2), gradient architectures (US-12283655-B2), electrode integration (US-12463249-B2), and manufacturing processes (US-20230035720-A1). Claude identified one Solid Energies patent (US 11,967,678) and correctly rated it as the highest-priority FTO concern but did not surface the broader portfolio. ChatGPT and Co-Pilot identified zero Solid Energies filings.
The practical significance is that a company relying on any individual patent hit would underestimate the scope of Solid Energies' IP position. The fencing strategy—covering the composition, the architecture, the electrode integration, and the manufacturing method—means that identifying a single design-around for one patent does not resolve the FTO exposure from the portfolio as a whole. This is the kind of strategic insight that requires seeing the full picture, which no general-purpose model delivered
2.4 Assignee Attribution Quality
ChatGPT's response included at least two instances of fabricated or unverifiable assignee attributions. For US 11,367,895 B1, the listed assignee was "Likely startup / defense contractor cluster." For US 2021/0202983 A1, the assignee was described as "Likely DOE / national lab ecosystem." In both cases, the model appears to have inferred the assignee from contextual patterns in its training data rather than retrieving the information from patent records.
In any operational IP workflow, assignee identity is foundational. It determines licensing strategy, litigation risk, and competitive positioning. A fabricated assignee is more dangerous than a missing one because it creates an illusion of completeness that discourages further investigation. An R&D team receiving this output might reasonably conclude that the landscape analysis is finished when it is not.
3. Structural Limitations of General-Purpose Models for Patent Intelligence
3.1 Training Data Is Not Patent Data
Large language models are trained on web-scraped text. Their knowledge of the patent record is derived from whatever fragments appeared in their training corpus: blog posts mentioning filings, news articles about litigation, snippets of Google Patents pages that were crawlable at the time of data collection. They do not have systematic, structured access to the USPTO database. They cannot query patent classification codes, parse claim language against a specific technology architecture, or verify whether a patent has been assigned, abandoned, or subjected to terminal disclaimer since their training data was collected.
This is not a limitation that improves with scale. A larger training corpus does not produce systematic patent coverage; it produces a larger but still arbitrary sampling of the patent record. The result is that general-purpose models will consistently surface well-known patents from heavily discussed assignees (QuantumScape, for example, appeared in most responses) while missing commercially significant filings from less publicly visible entities (Solid Energies, Korea Institute of EnergyResearch, Shenzhen Solid Advanced Materials).
3.2 The Web Is Closing to Model Scrapers
The data access problem is structural and worsening. As of mid-2025, Cloudflare reported that among the top 10,000 web domains, the majority now fully disallow AI crawlers such as GPTBot andClaudeBot via robots.txt. The trend has accelerated from partial restrictions to outright blocks, and the crawl-to-referral ratios reveal the underlying tension: OpenAI's crawlers access approximately1,700 pages for every referral they return to publishers; Anthropic's ratio exceeds 73,000 to 1.
Patent databases, scientific publishers, and IP analytics platforms are among the most restrictive content categories. A Duke University study in 2025 found that several categories of AI-related crawlers never request robots.txt files at all. The practical consequence is that the knowledge gap between what a general-purpose model "knows" about the patent landscape and what actually exists in the patent record is widening with each training cycle. A landscape query that a general-purpose model partially answered in 2023 may return less useful information in 2026.
3.3 General-Purpose Models Lack Ontological Frameworks for Patent Analysis
A freedom-to-operate analysis is not a summarization task. It requires understanding claim scope, prosecution history, continuation and divisional chains, assignee normalization (a single company may appear under multiple entity names across patent records), priority dates versus filing dates versus publication dates, and the relationship between dependent and independent claims. It requires mapping the specific technical features of a proposed product against independent claim language—not keyword matching.
General-purpose models do not have these frameworks. They pattern-match against training data and produce outputs that adopt the format and tone of patent analysis without the underlying data infrastructure. The format is correct. The confidence is high. The coverage is incomplete in ways that are not visible to the user.
4. Comparative Output Quality
The following table summarizes the qualitative characteristics of each tool's response across the dimensions most relevant to an operational IP workflow.

5. Implications for R&D and IP Organizations
5.1 The Confidence Problem
The central risk identified by this study is not that general-purpose models produce bad outputs—it is that they produce incomplete outputs with high confidence. Each model delivered its results in a professional format with structured analysis, risk ratings, and strategic recommendations. At no point did any model indicate the boundaries of its knowledge or flag that its results represented a fraction of the available patent record. A practitioner receiving one of these outputs would have no signal that the analysis was incomplete unless they independently validated it against a comprehensive datasource.
This creates an asymmetric risk profile: the better the format and tone of the output, the less likely the user is to question its completeness. In a corporate environment where AI outputs are increasingly treated as first-pass analysis, this dynamic incentivizes under-investigation at precisely the moment when thoroughness is most critical.
5.2 The Diversification Illusion
It might be assumed that running the same query through multiple general-purpose models provides validation through diversity of sources. This study suggests otherwise. While the four tools returned different subsets of patents, all operated under the same structural constraints: training data rather than live patent databases, web-scraped content rather than structured IP records, and general-purpose reasoning rather than patent-specific ontological frameworks. Running the same query through three constrained tools does not produce triangulation; it produces three partial views of the same incomplete picture.
5.3 The Appropriate Use Boundary
General-purpose language models are effective tools for a wide range of tasks: drafting communications, summarizing documents, generating code, and exploratory research. The finding of this study is not that these tools lack value but that their value boundary does not extend to decisions that carry existential commercial risk.
Patent landscape analysis, freedom-to-operate assessment, and competitive intelligence that informs R&D investment decisions fall outside that boundary. These are workflows where the completeness and verifiability of the underlying data are not merely desirable but are the primary determinant of whether the analysis has value. A patent landscape that captures 10% of the relevant filings, regardless of how well-formatted or confidently presented, is a liability rather than an asset.
6. Test 2: Competitive Intelligence — Bio-Based Polyamide Patent Landscape
To assess whether the findings from Test 1 were specific to a single technology domain or reflected a broader structural pattern, a second query was submitted to all four tools. This query shifted from freedom-to-operate analysis to competitive intelligence, asking each tool to identify the top 10organizations by patent filing volume in bio-based polyamide synthesis from castor oil derivatives over the past three years, with summaries of technical approach, co-assignee relationships, and portfolio trajectory.
6.1 Query

6.2 Summary of Results

6.3 Key Differentiators
Verifiability
The most consequential difference in Test 2 was the presence or absence of verifiable evidence. Cypris cited over 100 individual patent filings with full patent numbers, assignee names, and publication dates. Every claim about an organization’s technical focus, co-assignee relationships, and filing trajectory was anchored to specific documents that a practitioner could independently verify in USPTO, Espacenet, or WIPO PATENT SCOPE. No general-purpose model cited a single patent number. Claude produced the most structured and analytically useful output among the public models, with estimated filing ranges, product names, and strategic observations that were directionally plausible. However, without underlying patent citations, every claim in the response requires independent verification before it can inform a business decision. ChatGPT and Co-Pilot offered thinner profiles with no filing counts and no patent-level specificity.
Data Integrity
ChatGPT’s response contained a structural error that would mislead a practitioner: it listed CathayBiotech as organization #5 and then listed “Cathay Affiliate Cluster” as a separate organization at #9, effectively double-counting a single entity. It repeated this pattern with Toray at #4 and “Toray(Additional Programs)” at #10. In a competitive intelligence context where the ranking itself is the deliverable, this kind of error distorts the landscape and could lead to misallocation of competitive monitoring resources.
Organizations Missed
Cypris identified Kingfa Sci. & Tech. (8–10 filings with a differentiated furan diacid-based polyamide platform) and Zhejiang NHU (4–6 filings focused on continuous polymerization process technology)as emerging players that no general-purpose model surfaced. Both represent potential competitive threats or partnership opportunities that would be invisible to a team relying on public AI tools.Conversely, ChatGPT included organizations such as ANTA and Jiangsu Taiji that appear to be downstream users rather than significant patent filers in synthesis, suggesting the model was conflating commercial activity with IP activity.
Strategic Depth
Cypris’s cross-cutting observations identified a fundamental chemistry divergence in the landscape:European incumbents (Arkema, Evonik, EMS) rely on traditional castor oil pyrolysis to 11-aminoundecanoic acid or sebacic acid, while Chinese entrants (Cathay Biotech, Kingfa) are developing alternative bio-based routes through fermentation and furandicarboxylic acid chemistry.This represents a potential long-term disruption to the castor oil supply chain dependency thatWestern players have built their IP strategies around. Claude identified a similar theme at a higher level of abstraction. Neither ChatGPT nor Co-Pilot noted the divergence.
6.4 Test 2 Conclusion
Test 2 confirms that the coverage and verifiability gaps observed in Test 1 are not domain-specific.In a competitive intelligence context—where the deliverable is a ranked landscape of organizationalIP activity—the same structural limitations apply. General-purpose models can produce plausible-looking top-10 lists with reasonable organizational names, but they cannot anchor those lists to verifiable patent data, they cannot provide precise filing volumes, and they cannot identify emerging players whose patent activity is visible in structured databases but absent from the web-scraped content that general-purpose models rely on.
7. Conclusion
This comparative analysis, spanning two distinct technology domains and two distinct analytical workflows—freedom-to-operate assessment and competitive intelligence—demonstrates that the gap between purpose-built R&D intelligence platforms and general-purpose language models is not marginal, not domain-specific, and not transient. It is structural and consequential.
In Test 1 (LLZO garnet electrolytes for Li-S batteries), the purpose-built platform identified more than three times as many patents as the best-performing general-purpose model and ten times as many as the lowest-performing one. Among the patents identified exclusively by the purpose-built platform were filings rated as Very High FTO risk that directly claim the proposed technology architecture. InTest 2 (bio-based polyamide competitive landscape), the purpose-built platform cited over 100individual patent filings to substantiate its organizational rankings; no general-purpose model cited as ingle patent number.
The structural drivers of this gap—reliance on training data rather than live patent feeds, the accelerating closure of web content to AI scrapers, and the absence of patent-specific analytical frameworks—are not transient. They are inherent to the architecture of general-purpose models and will persist regardless of increases in model capability or training data volume.
For R&D and IP leaders, the practical implication is clear: general-purpose AI tools should be used for general-purpose tasks. Patent intelligence, competitive landscaping, and freedom-to-operate analysis require purpose-built systems with direct access to structured patent data, domain-specific analytical frameworks, and the ability to surface what a general-purpose model cannot—not because it chooses not to, but because it structurally cannot access the data.
The question for every organization making R&D investment decisions today is whether the tools informing those decisions have access to the evidence base those decisions require. This study suggests that for the majority of general-purpose AI tools currently in use, the answer is no.
Study Disclosure
This comparative evaluation was commissioned and published by Cypris. The testing methodology, prompts, evaluation criteria, and underlying outputs have been documented to support independent review and replication.
All platform outputs were preserved in their original form. Patent data and material factual claims were cross-checked against USPTO Patent Center and WIPO PATENTSCOPE records as of March 27, 2026. Cypris was one of the platforms evaluated and therefore has a commercial interest in the findings.
The Patent Intelligence Gap - A Comparative Analysis of Verticalized AI-Patent Tools vs. General-Purpose Language Models for R&D Decision-Making
Blogs

AAV gene therapy has crossed into commercial reality, and its patent landscape is distinctive because an AAV therapy is a modular product whose parts are patented separately. An adeno-associated virus vector delivers a therapeutic gene by packaging it inside an engineered protein shell, the capsid, whose surface engages target-cell receptors and whose fate through endocytosis, endosomal escape, and nuclear import determines where the therapy goes and how the immune system responds to it.¹ The intellectual property divides across distinct regions, each a distinct area of patenting: the capsid, spanning natural serotypes and, increasingly, engineered capsids produced by directed evolution, structure-guided design, and machine-learning-guided design;²,³,⁴,⁵ the strategies that address immunogenicity, because pre-existing neutralizing antibodies exclude many patients and the immune response generally prevents redosing;⁶,⁷ the transgene expression cassette, including the promoter and regulatory elements that control where and how strongly the gene is expressed; and the manufacturing process, where empty-capsid content, host-cell productivity, and the ratio of full to empty particles remain challenges at commercial scale.⁸ Because a therapy depends on several of these layers and they are often held by different owners, freedom-to-operate for an AAV product is a multi-layer, multi-owner analysis rather than a single clearance.
The competitive and legal environment has raised the stakes across every layer. Capsid engineering is the most active area of IP, with specialized platform companies developing next-generation capsids that target specific tissues such as the brain, muscle, and retina and that aim to evade pre-existing immunity, and licensing these platforms to larger developers.⁴,⁹ At the same time, foundational vector patents have been litigated, with the Federal Circuit deciding an appeal on core AAV vector claims in early 2026, so the boundaries of the foundational estate continue to be tested even as the field advances.¹⁰ The commercial ground is set by a broader wave of approved products: as of the 2024 review window, the US Food and Drug Administration had approved fourteen cellular and gene therapy products across the modality, of which the AAV-based Luxturna (voretigene neparvovec-rzyl) for RPE65-mediated inherited retinal dystrophy remains the AAV exemplar, first approved in the United States in 2017 and separately authorized in the European Union.⁴,¹¹,¹² Because applications publish about eighteen months after filing, the newest capsid and immune-evasion filings are under-represented, so the current frontier is more active than granted-patent counts suggest.
The strategic question is where defensible, hard-to-design-around IP sits. Across the Cypris corpus of more than 500 million patents and scientific papers, the AAV capsid-variant set holds on the order of 12,966 families and grew from about 828 in 2020 to roughly 1,818 in 2024, with the most active assignees led by the University of Pennsylvania, Voyager Therapeutics, the University of Massachusetts, Genzyme, and UC San Diego, and the United States far ahead of China, France, and the United Kingdom on geography; 2025 and 2026 counts are partial because of the publication lag. The most active and high-value ground is in engineered capsids that both target a tissue precisely and evade pre-existing immunity, because these directly address the field's central limitations, and receptor-guided and machine-learning-guided capsid design are advancing quickly here.²,³,⁴,⁵ Redosing and immune-modulation technologies are a distinct and comparatively open layer,⁶,⁷ as are manufacturing methods that raise the full-to-empty ratio and lower cost,⁸ and expression-cassette designs that improve durability and tissue specificity.⁴,⁹ Reading the landscape by layer and by owner, and tracking both the patents and the underlying virology and immunology research, is what separates a workable position from a blocked one.
What creates FTO risk in AAV gene therapy
Capsid claims. These cover natural serotypes and engineered capsids from directed evolution, structure-guided, and machine-learning design, the most active and contested layer.²,³,⁴,⁵
Immunogenicity and redosing claims. These cover strategies to evade pre-existing antibodies and enable redosing, a distinct and high-value layer given the field's central limitation.⁶,⁷
Transgene and expression-cassette claims. These cover promoters, regulatory elements, and the engineered transgene that control expression, a separately owned layer.
Manufacturing and purification claims. These cover producer systems, full-to-empty separation, host-cell productivity, and purification, where practical, hard-to-design-around barriers concentrate.⁸
Tissue-targeting claims. These cover receptor-guided and tissue-specific delivery, which can independently constrain a competing program.⁹
How AI-powered landscape and FTO analysis helps
A modular, multi-owner, litigation-shaped landscape is beyond manual clearance. AI-powered analysis addresses this with semantic search that retrieves relevant capsid, immunogenicity, cassette, and manufacturing claims regardless of terminology, attribution that resolves the many owners and license chains to canonical entities, claim-level analysis that separates the layers, and continuous monitoring that tracks new filings and disputes. Because AAV advances appear in virology and immunology literature before they are patented, reading both patents and literature gives earlier warning of where the field is heading.
Where Cypris fits
Cypris runs patent landscape and freedom-to-operate analysis for modular, contested fields such as AAV gene therapy across a corpus of more than 500 million patents and scientific papers, organized through a proprietary R&D ontology. The ontology clusters the landscape by layer, capsid, immunogenicity, expression cassette, and manufacturing, and normalizes developer, platform-specialist, and academic owners and their license chains to canonical entities, so a team sees how rights are distributed across the many parties rather than a flat list. Semantic search across patents and scientific literature surfaces relevant claims regardless of terminology and connects filings to the underlying research, which is where new capsids and immune-evasion strategies emerge first. Cypris Q, the platform's agentic layer, lets teams run landscape and FTO analysis conversationally and chain the attribution, clustering, and claim-level analysis across layers, and Agentic Monitoring tracks the landscape over time and flags new filings and developments as they publish. Cypris provides enterprise API partnerships with OpenAI, Anthropic, and Google, and is built with enterprise-grade security. Cypris serves hundreds of enterprise customers across pharmaceuticals, chemicals, advanced materials, energy, and other regulated industries.
FAQ
Why is freedom-to-operate hard for AAV gene therapy? Freedom-to-operate is hard for AAV gene therapy because a therapy is built from a capsid, an immune-evasion strategy, a transgene expression cassette, and a manufacturing process, each independently patentable and often held by different owners. The capsid layer alone is heavily engineered and contested. FTO must be assessed layer by layer across multiple estates.
Why is the capsid the key layer? The capsid is the key layer because it determines which tissues the therapy reaches and how the immune system responds, and it is where most engineering and patenting activity concentrates. Engineered capsids from directed evolution, structure-guided design, and machine learning aim to target tissues and evade immunity. That makes capsid IP the most active battleground.
What claim types create FTO risk in AAV therapy? Five claim types create FTO risk: capsid claims, immunogenicity and redosing claims, transgene and expression-cassette claims, manufacturing and purification claims, and tissue-targeting claims. Each covers a distinct layer and can be held by a different owner. The capsid and manufacturing layers are especially decisive.
Why do immunity and redosing matter so much? Immunity and redosing matter because pre-existing neutralizing antibodies exclude many patients from AAV therapy, and the immune response to a first dose generally prevents giving a second, so AAV is typically a single-dose modality. Technologies that evade pre-existing immunity or enable redosing address a central limitation. They are therefore a distinct, high-value layer.
Where is the white space in AAV gene therapy? The white space includes engineered capsids that target tissues and evade pre-existing immunity, redosing and immune-modulation technologies, manufacturing methods that raise the full-to-empty ratio, and expression-cassette designs that improve durability and specificity. Natural serotypes and liver-directed approaches are comparatively crowded. The durable, defensible value is in capsids, immune evasion, and manufacturing.
Why does AAV analysis need scientific literature? AAV analysis needs scientific literature because new capsids, immune-evasion strategies, and manufacturing advances appear in virology and immunology research before they are patented, so the literature gives the earliest signal. Analyzing patents alone gives a lagging view. Cypris analyzes both across more than 500 million patents and scientific papers.
What software helps analyze the AAV gene therapy patent landscape? Software for the AAV landscape should resolve developer, platform-specialist, and academic owners and license chains to canonical entities, cluster the capsid, immunogenicity, cassette, and manufacturing layers, search patents and scientific literature semantically, and monitor litigation and new filings continuously. Cypris does this across more than 500 million patents and scientific papers using a proprietary R&D ontology, semantic search, Cypris Q, and Agentic Monitoring.
Which teams need AAV patent landscape and FTO analysis? AAV patent landscape and FTO analysis is needed by R&D, IP, and business-development teams at gene therapy and pharmaceutical companies, capsid-platform specialists, academic technology-transfer offices, and investors assessing gene therapy assets. The modular, litigation-shaped landscape makes structured analysis essential. Cypris serves hundreds of enterprise customers across pharmaceuticals and other research-intensive industries.
Endnotes
- Hao, Y., & Xiang, J. (2023). Biophysical characterization of the AAV capsid through the viral transduction life cycle. Journal of Genetic Engineering and Biotechnology, 21. https://doi.org/10.1186/s43141-023-00518-5
- Fakhiri, J., Becker, S., & Grimm, D. (2022). Fantastic AAV gene therapy vectors and how to find them: random diversification, rational design and machine learning. Pathogens, 11(7), 756. https://doi.org/10.3390/pathogens11070756
- Qi, Y., et al. (2025). Artificial intelligence-based approaches for AAV vector engineering. Advanced Science, 12. https://doi.org/10.1002/advs.202411062
- Gao, F., et al. (2024). AAV engineering and load strategy for tropism modification, immune evasion and enhanced transgene expression. International Journal of Nanomedicine, 19. https://doi.org/10.2147/ijn.s459905
- Fu, W., et al. (2024). Machine-learning-guided directed evolution for AAV capsid engineering. Current Pharmaceutical Design. https://doi.org/10.2174/0113816128286593240226060318
- Barnes, C., Scheideler, O., & Schaffer, D. (2019). Engineering the AAV capsid to evade immune responses. Current Opinion in Biotechnology, 60. https://doi.org/10.1016/j.copbio.2019.01.002
- Meumann, N., Rodríguez-Márquez, E., & Büning, H. (2020). AAV capsid engineering in liver-directed gene therapy. Expert Opinion on Biological Therapy, 21(6). https://doi.org/10.1080/14712598.2021.1865303
- Yoon, S., Lee, D., et al. (2024). Decoding cellular mechanism of rAAV and engineering host-cell factories. Biotechnology Advances, 72. https://doi.org/10.1016/j.biotechadv.2024.108322
- Marković, I., et al. (2025). AAV gene therapy drug development and translation of engineered ocular and neurotropic capsids. Clinical and Translational Science, 18. https://doi.org/10.1111/cts.70428
- RegenxBio Inc. v. Sarepta Therapeutics, Inc., No. 24-1408 (Fed. Cir. Feb. 20, 2026).
- U.S. Food and Drug Administration. Approved cellular and gene therapy products. https://www.fda.gov/vaccines-blood-biologics/cellular-gene-therapy-products/approved-cellular-and-gene-therapy-products
- U.S. Food and Drug Administration (2017). Luxturna (voretigene neparvovec-rzyl). https://www.fda.gov/vaccines-blood-biologics/cellular-gene-therapy-products/luxturna

Chemical intelligence unifies three data types that chemistry R&D depends on: patents, scientific literature, and chemical structure data. A question about a compound, a reaction, or a material rarely lives in one of these alone. The relevant disclosure may sit in a patent claim, a journal paper, or a structure database, and the connection between them is where the insight is.
Most tools address only one layer. Structure databases index compounds, patent databases index filings, and literature databases index papers, and researchers toggle between them manually. That fragmentation is slow and lossy: a compound found in one system is not automatically linked to the patents that claim it or the papers that characterize it.
In 2026, AI-powered chemical intelligence closes that gap. Semantic search and a structured model of the field retrieve across patents, papers, and structures together. This article defines chemical intelligence, explains why siloed search falls short, and describes how the AI-powered approach works.
What chemical intelligence covers
Chemical intelligence spans the full evidence base for a compound or material. It includes patents and published applications, peer-reviewed papers and preprints, chemical compound and structure data, synthesis and reaction information, and regulatory and commercial signals. The defining feature is unification: the same compound is connected across every source in which it appears.
This is broader than chemical patent search. Patent search answers what has been filed; chemical intelligence answers what is known about a compound or material across the literature, the patent record, and structure data at once, which is what R&D and IP teams in chemistry, materials, and pharmaceuticals actually need.
Why siloed chemical search falls short
Siloed search forces a researcher to run the same question three times, in three systems, with three query languages, and then reconcile the results by hand. Connections are missed because no single tool sees all the evidence. A compound identified in a structure database is not tied to the patents that claim it or the papers that report its properties.
Keyword search compounds the problem. In chemistry, the same compound or reaction is described under different names, notations, and terminology, so a keyword query misses filings and papers that use unexpected language. The volume of new chemistry filings and publications continues to rise, widening the gap between what a manual, siloed search finds and what actually exists.
How AI-powered chemical intelligence works
AI-powered chemical intelligence applies semantic search across a unified corpus of patents and scientific literature, retrieving disclosures by meaning rather than exact terms. This surfaces the papers and filings that describe a compound or reaction in different language, which keyword search overlooks.
An R&D ontology links the layers. Because an ontology is a structured map of technical concepts and their relationships, it connects a compound to the patents that claim it, the papers that characterize it, and the technology domains it belongs to. That linkage is what turns three separate result sets into one coherent picture.
Agentic workflows then operate on that picture. On an AI-native platform such as Cypris, an agent can assess chemical freedom-to-operate at the claim level, assemble a competitive landscape of a chemical technology, or monitor a compound class continuously, retrieving across patents, papers, and structure data and returning cited output.
Where chemical intelligence is used
Chemical freedom-to-operate is a primary use. Chemical FTO assesses whether making, using, or selling a compound or formulation would infringe active patent claims, and it depends on retrieving claims that may describe the same chemistry in different terms. Competitive monitoring is another: teams track competitor chemical patents and pipelines continuously rather than rebuilding a picture each quarter.
Materials and formulation scouting is a third. Researchers use chemical intelligence to identify sustainable material alternatives, track new synthesis trends, and find who is active in a compound class, drawing on patents and literature together. Each of these questions is answered more completely when structure, patent, and literature evidence is unified.
Chemical intelligence in practice
Cypris is an AI-native R&D intelligence platform that unifies chemical evidence across a corpus of more than 500 million patents and scientific papers, organized through a proprietary R&D ontology, alongside chemical compound data. The ontology links compounds to the patents that claim them and the papers that characterize them, so semantic search retrieves across all of it rather than one silo.
Cypris Q, the platform's agentic layer, runs chemical FTO, landscape, and prior art workflows and returns cited output, while Agentic Monitoring tracks compound classes and competitor chemical activity continuously across patents, scientific literature, chemical compound data, and regulatory sources. Cypris operates under enterprise API partnerships with OpenAI, Anthropic, and Google, with enterprise-grade security, and serves hundreds of enterprise customers across pharmaceuticals, chemicals, advanced materials, and other regulated industries.
FAQ
What is a chemical intelligence platform?
A chemical intelligence platform unifies patents, scientific literature, and chemical structure data so R&D teams can search all three together rather than in separate silos. It connects a compound to the patents that claim it and the papers that characterize it, which is broader than chemical patent search alone.
What data does chemical intelligence cover?
Chemical intelligence covers patents and published applications, peer-reviewed papers and preprints, chemical compound and structure data, synthesis and reaction information, and regulatory and commercial signals. The defining feature is that the same compound is linked across every source in which it appears.
Can I search patents and chemical structures together?
Searching patents and chemical structures together requires a platform that unifies both in one corpus and links compounds to the filings that claim them. An AI-powered chemical intelligence platform does this with semantic search and an R&D ontology, so a compound and its patent coverage are connected rather than searched separately.
Is there a platform to search scientific papers and chemical structures?
A chemical intelligence platform searches scientific papers and chemical structures together by unifying literature and compound data in a single corpus. This matters because a compound's properties are often reported in papers before or alongside its appearance in patents, so searching both together gives a fuller picture.
How does AI improve chemical patent research?
AI improves chemical patent research by applying semantic search, which retrieves filings that describe the same compound or reaction in different names and notations. Combined with an R&D ontology that links compounds to their patents and papers, it surfaces evidence that keyword search across a single database misses.
What is chemical freedom-to-operate (FTO)?
Chemical freedom-to-operate assesses whether making, using, or selling a compound or formulation would infringe active patent claims. It depends on retrieving claims that may describe the same chemistry in different terms, which is why semantic search across a unified corpus is central to reliable chemical FTO.
How do R&D teams monitor competitor chemical patents?
R&D teams monitor competitor chemical patents most effectively with continuous, AI-powered monitoring that interprets new filings in the context of a compound class or technology domain. This replaces quarterly manual rebuilds and surfaces competitor chemical activity as it publishes.
Can chemical intelligence track new material synthesis trends?
Chemical intelligence can track new material synthesis trends by analyzing patents and scientific literature together and grouping activity by technical concept. This reveals where synthesis routes and material classes are developing, and which organizations are active, earlier than a patent-only view.
How does semantic search work for chemistry?
Semantic search for chemistry retrieves patents and papers by the meaning of a compound, reaction, or property rather than exact keywords. Because chemistry is described under many names and notations, semantic retrieval surfaces relevant disclosures that literal term matching overlooks.
What is the best chemical intelligence platform for R&D teams?
The best chemical intelligence platform unifies patents, scientific literature, and chemical structure data with semantic search and citable output. Cypris runs chemical intelligence on a corpus of more than 500 million patents and scientific papers organized through a proprietary R&D ontology, alongside chemical compound data, linking compounds to their patents and publications.

Prior art search determines whether an invention has already been disclosed publicly, anywhere, before a given date. It underpins patentability decisions, invalidity challenges, and R&D direction. If relevant prior art exists and is missed, a patent may be granted on shaky ground, or a competitor's patent may go unchallenged when it could have been invalidated.
Prior art is not limited to patents. It includes scientific papers, conference proceedings, technical disclosures, product documentation, and other public information. This is why prior art search must span patents and scientific literature together, and why patent-only searching leaves gaps, especially in fields where research is published before it is patented.
In 2026, AI-powered prior art search applies semantic search across a unified corpus of patents and scientific literature, retrieving conceptually relevant disclosures regardless of the exact words used. This article explains how it works and how to run one.
What counts as prior art
Prior art is any public disclosure of an invention before the relevant date. It includes granted patents and published applications, but also peer-reviewed papers, preprints, conference materials, theses, standards documents, and public product information. A disclosure in any of these can defeat novelty or support an obviousness argument.
Because prior art spans formats and languages, coverage and recall are the central challenges. A search that only covers patents, or only covers one language, systematically misses disclosures that exist elsewhere. The goal of prior art search is to find the most relevant disclosures, not simply to return many documents.
Prior art search versus freedom-to-operate
Prior art search and freedom-to-operate search are often confused because they use overlapping data, but they answer different questions. Prior art search asks whether an invention is new and non-obvious, which bears on whether a patent should be granted or can be invalidated. Freedom-to-operate search asks whether commercializing a product would infringe active, in-force patent claims.
The distinction changes what each search prioritizes. Prior art search values broad recall across patents and scientific literature to establish what was already known. FTO search focuses on active claims in specific jurisdictions to assess infringement risk. Using the right search for the question is essential to reaching a defensible conclusion.
How AI-powered prior art search works
AI-powered prior art search applies semantic search, which represents the meaning of text so that conceptually similar disclosures are retrieved even when the wording differs. This directly addresses the core weakness of keyword prior art search, where a relevant paper or patent is missed because it describes the invention in different terms.
Searching patents and scientific literature in a single unified corpus is what makes AI prior art search comprehensive. Early disclosure frequently appears in the literature before it reaches granted claims, particularly in biotech, chemistry, and materials science, so a unified search surfaces disclosures that a patent-only search cannot. An R&D ontology strengthens this by interpreting queries in the context of a technology domain, improving recall for the concepts that matter.
Agentic processes extend prior art search into an end-to-end workflow. An agent can expand a query into related concepts, retrieve candidate disclosures across patents and literature, summarize each with its relevance to the claims in question, and assemble a cited prior art report, with human experts reviewing and refining the result.
How to run an AI-powered prior art search
Begin by stating the invention and its key features precisely, and set the relevant date. Convert each feature into a semantic query so that conceptually equivalent disclosures are retrieved, not only exact-term matches. Run the search across a corpus that unifies patents and scientific literature, so that non-patent disclosures are captured.
Review candidate disclosures for relevance to the specific claims or features, and separate documents that anticipate the invention from those relevant to obviousness. For an invalidity search, map each strong reference to the claim elements it discloses. Assemble the findings into a cited report, and, where the position needs to stay current, place the technology area under continuous monitoring so that newly published disclosures are assessed as they appear.
Where Cypris fits
Cypris is an AI-native R&D intelligence platform that runs prior art search with semantic search across a corpus of more than 500 million patents and scientific papers, organized through a proprietary R&D ontology. The unified corpus and ontology let Cypris retrieve conceptually relevant disclosures across both patents and scientific literature, rather than matching keywords in patents alone.
Cypris Q, the agentic layer, expands queries, retrieves candidate disclosures, and assembles cited output, while Agentic Monitoring keeps a technology area current as new disclosures publish. Cypris operates under enterprise API partnerships with OpenAI, Anthropic, and Google, with enterprise-grade security, and serves hundreds of enterprise customers across pharmaceuticals, chemicals, advanced materials, and other regulated industries.
FAQ
What is a prior art search?
A prior art search determines whether an invention has already been disclosed publicly before a given date, across patents and non-patent sources. It underpins patentability decisions and invalidity challenges, because any earlier public disclosure can defeat novelty or support an obviousness argument.
What counts as prior art?
Prior art is any public disclosure of an invention before the relevant date, including granted patents, published applications, peer-reviewed papers, preprints, conference materials, theses, standards, and public product information. A disclosure in any of these formats can be relevant to novelty or obviousness.
What is the difference between prior art search and FTO?
Prior art search asks whether an invention is new and non-obvious, while freedom-to-operate search asks whether commercializing a product would infringe active patent claims. They use overlapping data but prioritize differently: prior art search values broad recall, and FTO focuses on active claims in specific jurisdictions.
Why must prior art search include scientific literature?
Prior art search must include scientific literature because early technical disclosure often appears in papers before it reaches granted patent claims, especially in biotech, chemistry, and materials science. A patent-only search systematically misses these non-patent disclosures.
How does AI improve prior art search?
AI improves prior art search by applying semantic search, which retrieves conceptually relevant disclosures even when the wording differs from the query. This addresses the main weakness of keyword prior art search, where relevant references are missed because they use unexpected terminology.
What is semantic prior art search?
Semantic prior art search represents the meaning of text so that conceptually similar disclosures are retrieved regardless of exact wording. It surfaces relevant patents and papers that keyword search overlooks, improving recall across a unified corpus of patents and scientific literature.
Can prior art search be automated with agents?
Prior art search can be automated with agentic processes that expand a query into related concepts, retrieve candidate disclosures across patents and literature, summarize each, and assemble a cited report. Human experts review and refine the output, while agents handle retrieval and synthesis at scale.
How do you run an invalidity prior art search?
An invalidity prior art search maps strong references to the specific claim elements they disclose, establishing what was already known before the relevant date. Semantic search across a unified corpus improves the chance of finding the anticipating or obviousness references that keyword search misses.
What data coverage does an effective prior art search need?
An effective prior art search needs broad coverage across patents and scientific literature in multiple languages, because prior art spans formats and jurisdictions. A corpus of more than 500 million patents and scientific papers organized through an R&D ontology supports the recall that prior art search requires.
What is the best software for prior art search?
The best prior art search software combines a unified corpus of patents and scientific literature with semantic search and citable output. Cypris runs prior art search across more than 500 million patents and scientific papers organized through a proprietary R&D ontology, retrieving conceptually relevant disclosures and assembling cited results.
Reports

This Cypris research brief maps the full ecosystem and value chain of electric vehicle battery systems and advanced battery materials, tracing the pathway from raw material extraction through precursor and active material production, cell component manufacturing, battery cell production, pack assembly, vehicle integration, and end-of-life recycling. The brief defines each segment's functional role, identifies key players across upstream, midstream, and downstream layers, and analyzes the structural forces — including critical mineral supply volatility, geographic concentration, OEM vertical integration strategies, recycling-driven circularity, and solid-state battery development — that are reshaping where value concentrates and where supply-chain risk resides.

This Cypris research brief maps the ecosystem and value chain of the specialty polymers and high-performance materials industry, covering the full pathway from raw material and monomer suppliers through polymer manufacturers, compounders, additive suppliers, specialty distributors, converters, and end-use OEMs across aerospace, automotive, electronics, medical, energy, and industrial markets. Beyond the segment-by-segment breakdown and player landscape, the brief analyzes the structural forces shaping the ecosystem — including vertical integration strategies, supplier concentration and consolidation patterns, geographic clustering, circularity constraints, and shifting end-market demand — with a central thesis that leverage in this ecosystem concentrates wherever technical specialization overlaps with requalification burden.

Cypris Research Services' inaugural Innovation Outlook examines how AI-driven data center demand is reshaping U.S. power infrastructure — and why hyperscalers have stopped waiting for the grid to catch up. The report synthesizes commercial activity, market sizing, technology trends, and patent-based competitive positioning into a single ecosystem view of behind-the-meter generation, sizing the U.S. opportunity at $35.8B and tracking 56 GW of contracted bypass capacity already in the pipeline. It identifies where the defensible whitespace actually sits — and it's not where most of the market is currently looking.
Webinars

Many enterprises have adopted horizontal, foundation-model AI platforms. But access to the same underlying models does not, by itself, create differentiated intelligence. For highly technical and mission-critical research, general-purpose models may produce broad but weakly grounded answers when they lack access to authoritative technical data, specialized context, and verifiable sources.
The next competitive advantage will come from the intelligence layer surrounding the foundation model: the domain-specific data, ontologies, retrieval capabilities, agent workflows, and source grounding that together form an AI harness. These verticalized systems can transform general-purpose AI into a more specialized capability for research, innovation, and technical decision-making.
Join Steve Hafif, Co-Founder and CEO of Cypris.ai, and Marlene Valderrama, Principal IP Manager and Senior Technology Scout at Halliburton, for a conversation on the state of enterprise AI and how organizations can enhance horizontal AI platforms with verticalized intelligence designed for R&D and innovation.
.png)

Most IP organizations are making high-stakes capital allocation decisions with incomplete visibility – relying primarily on patent data as a proxy for innovation. That approach is not optimal. Patents alone cannot reveal technology trajectories, capital flows, or commercial viability.
A more effective model requires integrating patents with scientific literature, grant funding, market activity, and competitive intelligence. This means that for a complete picture, IP and R&D teams need infrastructure that connects fragmented data into a unified, decision-ready intelligence layer.
AI is accelerating that shift. The value is no longer simply in retrieving documents faster; it’s in extracting signal from noise. Modern AI systems can contextualize disparate datasets, identify patterns, and generate strategic narratives – transforming raw information into actionable insight.
Join us on Thursday, April 23, at 12 PM ET for a discussion on how unified AI platforms are redefining decision-making across IP and R&D teams. Moderated by Gene Quinn, panelists Marlene Valderrama and Amir Achourie will examine how integrating technical, scientific, and market data collapses traditional silos – enabling more aligned strategy, sharper investment decisions, and measurable business impact.
Register here: https://ipwatchdog.com/cypris-april-23-2026/
.png)
In this session, we break down how AI is reshaping the R&D lifecycle, from faster discovery to more informed decision-making. See how an intelligence layer approach enables teams to move beyond fragmented tools toward a unified, scalable system for innovation.
.avif)
