
Resources
Guides, research, and perspectives on R&D intelligence, IP strategy, and the future of AI enabled innovation.

Executive Summary
In 2024, US patent infringement jury verdicts totaled $4.19 billion across 72 cases. Twelve individual verdicts exceeded $100million. The largest single award—$857 million in General Access Solutions v.Cellco Partnership (Verizon)—exceeded the annual R&D budget of many mid-market technology companies. In the first half of 2025 alone, total damages reached an additional $1.91 billion.
The consequences of incomplete patent intelligence are not abstract. In what has become one of the most instructive IP disputes in recent history, Masimo’s pulse oximetry patents triggered a US import ban on certain Apple Watch models, forcing Apple to disable its blood oxygen feature across an entire product line, halt domestic sales of affected models, invest in a hardware redesign, and ultimately face a $634 million jury verdict in November 2025. Apple—a company with one of the most sophisticated intellectual property organizations on earth—spent years in litigation over technology it might have designed around during development.
For organizations with fewer resources than Apple, the risk calculus is starker. A mid-size materials company, a university spinout, or a defense contractor developing next-generation battery technology cannot absorb a nine-figure verdict or a multi-year injunction. For these organizations, the patent landscape analysis conducted during the development phase is the primary risk mitigation mechanism. The quality of that analysis is not a matter of convenience. It is a matter of survival.
And yet, a growing number of R&D and IP teams are conducting that analysis using general-purpose AI tools—ChatGPT, Claude, Microsoft Co-Pilot—that were never designed for patent intelligence and are structurally incapable of delivering it.
This report presents the findings of a controlled comparison study in which identical patent landscape queries were submitted to four AI-powered tools: Cypris (a purpose-built R&D intelligence platform),ChatGPT (OpenAI), Claude (Anthropic), and Microsoft Co-Pilot. Two technology domains were tested: solid-state lithium-sulfur battery electrolytes using garnet-type LLZO ceramic materials (freedom-to-operate analysis), and bio-based polyamide synthesis from castor oil derivatives (competitive intelligence).
The results reveal a significant and structurally persistent gap. In Test 1, Cypris identified over 40 active US patents and published applications with granular FTO risk assessments. Claude identified 12. ChatGPT identified 7, several with fabricated attribution. Co-Pilot identified 4. Among the patents surfaced exclusively by Cypris were filings rated as “Very High” FTO risk that directly claim the technology architecture described in the query. In Test 2, Cypris cited over 100 individual patent filings with full attribution to substantiate its competitive landscape rankings. No general-purpose model cited a single patent number.
The most active sectors for patent enforcement—semiconductors, AI, biopharma, and advanced materials—are the same sectors where R&D teams are most likely to adopt AI tools for intelligence workflows. The findings of this report have direct implications for any organization using general-purpose AI to inform patent strategy, competitive intelligence, or R&D investment decisions.

1. Methodology
A controlled comparative evaluation was conducted on March 27, 2026. An identical patent landscape query was submitted verbatim to each platform under standardized testing conditions. No follow-up prompts, clarifications, or iterative refinements were permitted, ensuring that each platform was evaluated based solely on its initial response.
The outputs were preserved in their original form and evaluated against predefined criteria using publicly verifiable patent records.
1.1 Query
Identify all active US patents and published applications filed in the last 5 years related to solid-state lithium-sulfur battery electrolytes using garnet-type ceramic materials. For each, provide the assignee, filing date, key claims, and current legal status. Highlight any patents that could pose freedom-to-operate risks for a company developing a Li₇La₃Zr₂O₁₂(LLZO)-based composite electrolyte with a polymer interlayer.
1.2 Tools Evaluated

1.3 Evaluation Criteria
Each response was evaluated using a consistent six-part scoring framework: patent coverage, assignee accuracy, filing metadata completeness, depth of claim analysis, quality of FTO risk stratification, and the presence of actionable strategic guidance.
Patent numbers, assignees, filing information, and legal status were independently checked against publicly available USPTO and WIPO records. The evaluation focused on the completeness, accuracy, and practical utility of each platform’s output rather than writing quality or presentation.
2. Findings
2.1 Coverage Gap
The most significant finding is the scale of the coverage differential. Cypris identified over 40 active US patents and published applications spanning LLZO-polymer composite electrolytes, garnet interface modification, polymer interlayer architectures, lithium-sulfur specific filings, and adjacent ceramic composite patents. The results were organized by technology category with per-patent FTO risk ratings.
Claude identified 12 patents organized in a four-tier risk framework. Its analysis was structurally sound and correctly flagged the two highest-risk filings (Solid Energies US 11,967,678 and the LLZO nanofiber multilayer US 11,923,501). It also identified the University ofMaryland/ Wachsman portfolio as a concentration risk and noted the NASA SABERS portfolio as a licensing opportunity. However, it missed the majority of the landscape, including the entire Corning portfolio, GM's interlayer patents, theKorea Institute of Energy Research three-layer architecture, and the HonHai/SolidEdge lithium-sulfur specific filing.
ChatGPT identified 7 patents, but the quality of attribution was inconsistent. It listed assignees as "Likely DOE /national lab ecosystem" and "Likely startup / defense contractor cluster" for two filings—language that indicates the model was inferring rather than retrieving assignee data. In a freedom-to-operate context, an unverified assignee attribution is functionally equivalent to no attribution, as it cannot support a licensing inquiry or risk assessment.
Co-Pilot identified 4 US patents. Its output was the most limited in scope, missing the Solid Energies portfolio entirely, theUMD/ Wachsman portfolio, Gelion/ Johnson Matthey, NASA SABERS, and all Li-S specific LLZO filings.
2.2 Critical Patents Missed by Public Models
The following table presents patents identified exclusively by Cypris that were rated as High or Very High FTO risk for the proposed technology architecture. None were surfaced by any general-purpose model.

2.3 Patent Fencing: The Solid Energies Portfolio
Cypris identified a coordinated patent fencing strategy by Solid Energies, Inc. that no general-purpose model detected at scale. Solid Energies holds at least four granted US patents and one published application covering LLZO-polymer composite electrolytes across compositions(US-12463245-B2), gradient architectures (US-12283655-B2), electrode integration (US-12463249-B2), and manufacturing processes (US-20230035720-A1). Claude identified one Solid Energies patent (US 11,967,678) and correctly rated it as the highest-priority FTO concern but did not surface the broader portfolio. ChatGPT and Co-Pilot identified zero Solid Energies filings.
The practical significance is that a company relying on any individual patent hit would underestimate the scope of Solid Energies' IP position. The fencing strategy—covering the composition, the architecture, the electrode integration, and the manufacturing method—means that identifying a single design-around for one patent does not resolve the FTO exposure from the portfolio as a whole. This is the kind of strategic insight that requires seeing the full picture, which no general-purpose model delivered
2.4 Assignee Attribution Quality
ChatGPT's response included at least two instances of fabricated or unverifiable assignee attributions. For US 11,367,895 B1, the listed assignee was "Likely startup / defense contractor cluster." For US 2021/0202983 A1, the assignee was described as "Likely DOE / national lab ecosystem." In both cases, the model appears to have inferred the assignee from contextual patterns in its training data rather than retrieving the information from patent records.
In any operational IP workflow, assignee identity is foundational. It determines licensing strategy, litigation risk, and competitive positioning. A fabricated assignee is more dangerous than a missing one because it creates an illusion of completeness that discourages further investigation. An R&D team receiving this output might reasonably conclude that the landscape analysis is finished when it is not.
3. Structural Limitations of General-Purpose Models for Patent Intelligence
3.1 Training Data Is Not Patent Data
Large language models are trained on web-scraped text. Their knowledge of the patent record is derived from whatever fragments appeared in their training corpus: blog posts mentioning filings, news articles about litigation, snippets of Google Patents pages that were crawlable at the time of data collection. They do not have systematic, structured access to the USPTO database. They cannot query patent classification codes, parse claim language against a specific technology architecture, or verify whether a patent has been assigned, abandoned, or subjected to terminal disclaimer since their training data was collected.
This is not a limitation that improves with scale. A larger training corpus does not produce systematic patent coverage; it produces a larger but still arbitrary sampling of the patent record. The result is that general-purpose models will consistently surface well-known patents from heavily discussed assignees (QuantumScape, for example, appeared in most responses) while missing commercially significant filings from less publicly visible entities (Solid Energies, Korea Institute of EnergyResearch, Shenzhen Solid Advanced Materials).
3.2 The Web Is Closing to Model Scrapers
The data access problem is structural and worsening. As of mid-2025, Cloudflare reported that among the top 10,000 web domains, the majority now fully disallow AI crawlers such as GPTBot andClaudeBot via robots.txt. The trend has accelerated from partial restrictions to outright blocks, and the crawl-to-referral ratios reveal the underlying tension: OpenAI's crawlers access approximately1,700 pages for every referral they return to publishers; Anthropic's ratio exceeds 73,000 to 1.
Patent databases, scientific publishers, and IP analytics platforms are among the most restrictive content categories. A Duke University study in 2025 found that several categories of AI-related crawlers never request robots.txt files at all. The practical consequence is that the knowledge gap between what a general-purpose model "knows" about the patent landscape and what actually exists in the patent record is widening with each training cycle. A landscape query that a general-purpose model partially answered in 2023 may return less useful information in 2026.
3.3 General-Purpose Models Lack Ontological Frameworks for Patent Analysis
A freedom-to-operate analysis is not a summarization task. It requires understanding claim scope, prosecution history, continuation and divisional chains, assignee normalization (a single company may appear under multiple entity names across patent records), priority dates versus filing dates versus publication dates, and the relationship between dependent and independent claims. It requires mapping the specific technical features of a proposed product against independent claim language—not keyword matching.
General-purpose models do not have these frameworks. They pattern-match against training data and produce outputs that adopt the format and tone of patent analysis without the underlying data infrastructure. The format is correct. The confidence is high. The coverage is incomplete in ways that are not visible to the user.
4. Comparative Output Quality
The following table summarizes the qualitative characteristics of each tool's response across the dimensions most relevant to an operational IP workflow.

5. Implications for R&D and IP Organizations
5.1 The Confidence Problem
The central risk identified by this study is not that general-purpose models produce bad outputs—it is that they produce incomplete outputs with high confidence. Each model delivered its results in a professional format with structured analysis, risk ratings, and strategic recommendations. At no point did any model indicate the boundaries of its knowledge or flag that its results represented a fraction of the available patent record. A practitioner receiving one of these outputs would have no signal that the analysis was incomplete unless they independently validated it against a comprehensive datasource.
This creates an asymmetric risk profile: the better the format and tone of the output, the less likely the user is to question its completeness. In a corporate environment where AI outputs are increasingly treated as first-pass analysis, this dynamic incentivizes under-investigation at precisely the moment when thoroughness is most critical.
5.2 The Diversification Illusion
It might be assumed that running the same query through multiple general-purpose models provides validation through diversity of sources. This study suggests otherwise. While the four tools returned different subsets of patents, all operated under the same structural constraints: training data rather than live patent databases, web-scraped content rather than structured IP records, and general-purpose reasoning rather than patent-specific ontological frameworks. Running the same query through three constrained tools does not produce triangulation; it produces three partial views of the same incomplete picture.
5.3 The Appropriate Use Boundary
General-purpose language models are effective tools for a wide range of tasks: drafting communications, summarizing documents, generating code, and exploratory research. The finding of this study is not that these tools lack value but that their value boundary does not extend to decisions that carry existential commercial risk.
Patent landscape analysis, freedom-to-operate assessment, and competitive intelligence that informs R&D investment decisions fall outside that boundary. These are workflows where the completeness and verifiability of the underlying data are not merely desirable but are the primary determinant of whether the analysis has value. A patent landscape that captures 10% of the relevant filings, regardless of how well-formatted or confidently presented, is a liability rather than an asset.
6. Test 2: Competitive Intelligence — Bio-Based Polyamide Patent Landscape
To assess whether the findings from Test 1 were specific to a single technology domain or reflected a broader structural pattern, a second query was submitted to all four tools. This query shifted from freedom-to-operate analysis to competitive intelligence, asking each tool to identify the top 10organizations by patent filing volume in bio-based polyamide synthesis from castor oil derivatives over the past three years, with summaries of technical approach, co-assignee relationships, and portfolio trajectory.
6.1 Query

6.2 Summary of Results

6.3 Key Differentiators
Verifiability
The most consequential difference in Test 2 was the presence or absence of verifiable evidence. Cypris cited over 100 individual patent filings with full patent numbers, assignee names, and publication dates. Every claim about an organization’s technical focus, co-assignee relationships, and filing trajectory was anchored to specific documents that a practitioner could independently verify in USPTO, Espacenet, or WIPO PATENT SCOPE. No general-purpose model cited a single patent number. Claude produced the most structured and analytically useful output among the public models, with estimated filing ranges, product names, and strategic observations that were directionally plausible. However, without underlying patent citations, every claim in the response requires independent verification before it can inform a business decision. ChatGPT and Co-Pilot offered thinner profiles with no filing counts and no patent-level specificity.
Data Integrity
ChatGPT’s response contained a structural error that would mislead a practitioner: it listed CathayBiotech as organization #5 and then listed “Cathay Affiliate Cluster” as a separate organization at #9, effectively double-counting a single entity. It repeated this pattern with Toray at #4 and “Toray(Additional Programs)” at #10. In a competitive intelligence context where the ranking itself is the deliverable, this kind of error distorts the landscape and could lead to misallocation of competitive monitoring resources.
Organizations Missed
Cypris identified Kingfa Sci. & Tech. (8–10 filings with a differentiated furan diacid-based polyamide platform) and Zhejiang NHU (4–6 filings focused on continuous polymerization process technology)as emerging players that no general-purpose model surfaced. Both represent potential competitive threats or partnership opportunities that would be invisible to a team relying on public AI tools.Conversely, ChatGPT included organizations such as ANTA and Jiangsu Taiji that appear to be downstream users rather than significant patent filers in synthesis, suggesting the model was conflating commercial activity with IP activity.
Strategic Depth
Cypris’s cross-cutting observations identified a fundamental chemistry divergence in the landscape:European incumbents (Arkema, Evonik, EMS) rely on traditional castor oil pyrolysis to 11-aminoundecanoic acid or sebacic acid, while Chinese entrants (Cathay Biotech, Kingfa) are developing alternative bio-based routes through fermentation and furandicarboxylic acid chemistry.This represents a potential long-term disruption to the castor oil supply chain dependency thatWestern players have built their IP strategies around. Claude identified a similar theme at a higher level of abstraction. Neither ChatGPT nor Co-Pilot noted the divergence.
6.4 Test 2 Conclusion
Test 2 confirms that the coverage and verifiability gaps observed in Test 1 are not domain-specific.In a competitive intelligence context—where the deliverable is a ranked landscape of organizationalIP activity—the same structural limitations apply. General-purpose models can produce plausible-looking top-10 lists with reasonable organizational names, but they cannot anchor those lists to verifiable patent data, they cannot provide precise filing volumes, and they cannot identify emerging players whose patent activity is visible in structured databases but absent from the web-scraped content that general-purpose models rely on.
7. Conclusion
This comparative analysis, spanning two distinct technology domains and two distinct analytical workflows—freedom-to-operate assessment and competitive intelligence—demonstrates that the gap between purpose-built R&D intelligence platforms and general-purpose language models is not marginal, not domain-specific, and not transient. It is structural and consequential.
In Test 1 (LLZO garnet electrolytes for Li-S batteries), the purpose-built platform identified more than three times as many patents as the best-performing general-purpose model and ten times as many as the lowest-performing one. Among the patents identified exclusively by the purpose-built platform were filings rated as Very High FTO risk that directly claim the proposed technology architecture. InTest 2 (bio-based polyamide competitive landscape), the purpose-built platform cited over 100individual patent filings to substantiate its organizational rankings; no general-purpose model cited as ingle patent number.
The structural drivers of this gap—reliance on training data rather than live patent feeds, the accelerating closure of web content to AI scrapers, and the absence of patent-specific analytical frameworks—are not transient. They are inherent to the architecture of general-purpose models and will persist regardless of increases in model capability or training data volume.
For R&D and IP leaders, the practical implication is clear: general-purpose AI tools should be used for general-purpose tasks. Patent intelligence, competitive landscaping, and freedom-to-operate analysis require purpose-built systems with direct access to structured patent data, domain-specific analytical frameworks, and the ability to surface what a general-purpose model cannot—not because it chooses not to, but because it structurally cannot access the data.
The question for every organization making R&D investment decisions today is whether the tools informing those decisions have access to the evidence base those decisions require. This study suggests that for the majority of general-purpose AI tools currently in use, the answer is no.
Study Disclosure
This comparative evaluation was commissioned and published by Cypris. The testing methodology, prompts, evaluation criteria, and underlying outputs have been documented to support independent review and replication.
All platform outputs were preserved in their original form. Patent data and material factual claims were cross-checked against USPTO Patent Center and WIPO PATENTSCOPE records as of March 27, 2026. Cypris was one of the platforms evaluated and therefore has a commercial interest in the findings.
The Patent Intelligence Gap - A Comparative Analysis of Verticalized AI-Patent Tools vs. General-Purpose Language Models for R&D Decision-Making
Blogs

Artificial intelligence has become a permanent layer in pharmaceutical R&D, and it is generating a distinctive, fast-growing patent landscape. The convergence of AI and drug discovery is visible directly in the data on generative-AI patenting: among the categories tracked by the World Intellectual Property Organization, applications in molecules, genes, and proteins, though smaller in absolute number at roughly 1,500 inventions, were the fastest-growing, expanding at about 78 percent per year over a five-year period.¹ Peer-reviewed patent-basis analysis of AI in the pharmaceutical industry confirms both the rapid rise of filing activity and its concentration among a set of key players,² and a growing body of work applies patent-landscaping and bibliometric methods specifically to AI-driven drug discovery, including in areas such as cancer drug discovery.⁵,⁶,⁷ For R&D and IP teams, the strategic questions are which sub-domains are crowded, where defensible white space remains, and how the unsettled rules on AI-assisted inventorship affect what can be protected.
The landscape divides into several technically distinct sub-domains, each a different region of patenting. Generative molecular design covers models that propose novel candidate molecules. Drug-target interaction and binding-affinity prediction covers models that predict whether and how strongly a molecule binds a target. Drug repurposing covers methods that use biomedical knowledge graphs and network pharmacology to find new uses for known compounds. Multi-omics response prediction covers models that predict biological response from genomic and other omics data. Clinical-trial prediction and design covers models that forecast trial success and optimize design. And a further layer covers AI-assisted pharmaceutical development and, increasingly, generative AI applied to regulatory documentation. These sub-domains differ sharply in how crowded they are: biomedical knowledge-graph construction and traversal, for example, has become a comparatively crowded area of prior art, while newer large-language-model-native and agentic approaches are earlier and sparser.
Two structural features shape the landscape. The first is geographic and institutional concentration: the same concentration seen across generative AI, where a small number of countries account for most filings, extends into AI drug discovery, with China's share of generative-AI patenting near the top globally.¹ The second is the unsettled status of AI-assisted inventorship. In the United States, the Patent and Trademark Office rescinded its February 2024 guidance on AI-assisted inventions in November 2025 and returned to the traditional human-conception standard, which affects how AI-heavy pipelines document invention and how their patents should be valued.³ Commentators expect the first wave of litigation over AI-generated drug inventions within a few years, which will set precedents on inventorship and eligibility.⁴ These are not peripheral legal details; they determine what portion of an AI-driven discovery effort can be protected and how a portfolio should be structured, and they vary by jurisdiction. Because applications publish about eighteen months after filing, the newest large-language-model-native and agentic filings are under-represented, so the current frontier is more active than granted-patent counts suggest.
What the AI drug discovery landscape shows
Fastest-growing generative-AI category. Among generative-AI patents, molecule, gene, and protein applications grew fastest, at roughly 78 percent per year, though from a smaller base than image or text applications.¹
Several distinct sub-domains. The field spans generative molecular design, drug-target interaction prediction, knowledge-graph-based repurposing, multi-omics response prediction, clinical-trial prediction, and AI-assisted development.²
Crowded versus sparse areas. Biomedical knowledge-graph construction and traversal is comparatively crowded prior art, while large-language-model-native and agentic approaches are earlier and sparser.
Geographic concentration. Activity is concentrated in a small number of countries, mirroring the broader generative-AI landscape, with China prominent.¹
Unsettled inventorship. AI-assisted inventorship rules are in flux, with the US returning to a human-conception standard in late 2025, which affects what can be protected and how portfolios are documented.³,⁴
How AI-powered landscape and white space analysis helps
Mapping a fast-moving, sub-domain-structured field where much of the state of the art is in non-patent literature requires more than keyword search. AI-powered analysis addresses this with semantic search across both patents and scientific literature, which is essential because AI-method disclosures often appear first in preprints and conference proceedings, attribution that normalizes filers to canonical entities, and continuous monitoring that tracks the newest agentic and large-language-model-native filings. Clustering activity by sub-domain and by concept is what distinguishes crowded prior art from genuine white space.
Where Cypris fits
Cypris runs patent landscape and white space analysis for fast-moving, literature-heavy fields such as AI in drug discovery across a corpus of more than 500 million patents and scientific papers, organized through a proprietary R&D ontology. Because the corpus spans both patents and scientific literature, Cypris covers the preprints and conference proceedings where AI-method disclosures often appear first, rather than patents alone. The ontology clusters activity by sub-domain, generative design, interaction prediction, repurposing, multi-omics, and clinical prediction, and normalizes filers to canonical entities, so a team can resolve which areas, such as knowledge-graph methods, are crowded and which, such as agentic approaches, remain open as white space. Cypris Q, the platform's agentic layer, lets teams run landscape and white space analysis conversationally and chain the clustering, attribution, and gap analysis, and Agentic Monitoring tracks a defined area over time and flags new patents and papers as they publish, which is essential where the newest filings are under-represented by publication lag. Cypris provides enterprise API partnerships with OpenAI, Anthropic, and Google, and is built with enterprise-grade security. Cypris serves hundreds of enterprise customers across pharmaceuticals, chemicals, advanced materials, energy, and other regulated industries.
FAQ
How fast is AI drug discovery patenting growing?
AI drug discovery patenting is growing quickly. Among generative-AI patent categories tracked by WIPO, applications in molecules, genes, and proteins were the fastest-growing at roughly 78 percent per year, though from a smaller base than image or text applications. Peer-reviewed patent-basis analysis confirms the rapid rise of AI filing activity in the pharmaceutical industry.
What are the main sub-domains of AI drug discovery patents?
The main sub-domains are generative molecular design, drug-target interaction and binding-affinity prediction, drug repurposing using biomedical knowledge graphs, multi-omics response prediction, clinical-trial prediction and design, and AI-assisted pharmaceutical development. Each is a technically distinct region of the patent landscape. They differ substantially in how crowded they are.
Which areas of AI drug discovery are crowded, and which are open?
Biomedical knowledge-graph construction and traversal has become a comparatively crowded area of prior art in AI drug discovery, while large-language-model-native and agentic approaches are earlier and sparser. The crowded areas carry more freedom-to-operate risk, and the sparser areas hold more white space. Distinguishing them requires clustering activity by sub-domain and concept.
How does AI-assisted inventorship affect drug patents?
AI-assisted inventorship affects drug patents because the rules on whether and how AI-assisted inventions can be protected are unsettled and vary by jurisdiction. In the United States, the Patent and Trademark Office rescinded its 2024 guidance on AI-assisted inventions in November 2025 and returned to the traditional human-conception standard. This affects how AI-heavy pipelines document invention and how their patents are valued.
Why does China feature prominently in AI drug discovery patenting?
China features prominently in AI drug discovery patenting because it accounts for a large share of generative-AI patenting overall, and that concentration extends into the drug discovery sub-domains. The broader generative-AI landscape is dominated by a small number of countries. This geographic concentration matters for competitive positioning and freedom-to-operate.
Why does AI drug discovery analysis need scientific literature?
AI drug discovery analysis needs scientific literature because AI-method disclosures often appear first in preprints and conference proceedings rather than patents, so a patent-only view misses much of the state of the art. This is characteristic of AI fields generally. Cypris analyzes both patents and scientific literature across more than 500 million documents.
How do you find white space in AI drug discovery?
Finding white space in AI drug discovery means clustering activity by sub-domain and concept across patents and scientific literature, and identifying the sparser areas, such as agentic and large-language-model-native approaches, where few patents yet exist. Because much of the state of the art is in non-patent literature, semantic search across both sources is essential. The white space is where a viable method exists but patenting is still thin.
Which teams use AI drug discovery patent landscape analysis?
AI drug discovery patent landscape analysis is used by R&D, IP, and strategy teams at pharmaceutical companies, AI-native drug discovery firms, and their partners, as well as investors assessing AI-driven pipelines. It informs where to file, where freedom-to-operate risk sits, and how to structure a portfolio given inventorship uncertainty. Cypris serves hundreds of enterprise customers across pharmaceuticals and other research-intensive industries.
How current does an AI drug discovery landscape need to be?
An AI drug discovery landscape needs to be continuously current, because the field moves quickly, inventorship rules are shifting, new agentic and large-language-model-native filings publish constantly, and publication lag hides the most recent activity. A one-time landscape ages within months. Cypris uses Agentic Monitoring to track a defined area and flag new patents and papers as they publish.
Endnotes
- World Intellectual Property Organization (2024). Patent Landscape Report: Generative Artificial Intelligence. Geneva: WIPO. https://doi.org/10.34667/tind.49740
- Kano, S. & Sakaoka, S. (2025). Quantitative insights on artificial intelligence in the pharmaceutical industry: a patent-basis analysis of technological trends and key players. World Patent Information. https://www.sciencedirect.com/science/article/pii/S0172219025000481
- United States Patent and Trademark Office (2025). Revised Inventorship Guidance for AI-Assisted Inventions, Federal Register (published November 28, 2025; rescinding the February 2024 guidance and returning to the traditional human-conception standard). https://www.federalregister.gov/documents/2025/11/28/2025-21457/revised-inventorship-guidance-for-ai-assisted-inventions
- Goodwin (2026). AI Drug Discovery Tests the Limits of Patent Law. https://www.goodwinlaw.com/en/insights/publications/2025/12/insights-lifesciences-ip-ai-drug-discovery-tests-the-limits-of-patent-law
- Hofmann-Apitius, M., Gadiya, Y., Zaliani, A. & Gribbon, P. (2023). Pharmaceutical patent landscaping: a novel approach to understand patents from the drug discovery perspective. Artificial Intelligence in the Life Sciences. https://doi.org/10.1016/j.ailsci.2023.100061
- Abdulwahab, A. A. et al. (2024). Catalyzing innovation in cancer drug discovery through artificial intelligence, machine learning and patency. Pharmaceutical Patent Analyst.
- Jing, F. & Ma, Y. (2024). Bibliometric Analysis and Research Trends in Artificial Intelligence for Pharmaceutical Management and Drug Discovery.

The CRISPR and gene-editing patent landscape is one of the largest and most contested in biotechnology, and freedom-to-operate in this field is correspondingly difficult. A landscape analysis maintained by a national patent office counted roughly 23,700 CRISPR patent families as of the end of 2024, an increase of more than 6,500 families in a single year, and it identified four competing groups holding foundational claims.¹ Freedom-to-operate determines whether making, using, or selling a product would infringe another party's active patent claims. In gene editing, the foundational rights are split across multiple owners and jurisdictions, so a developer frequently cannot clear a product by licensing from a single source and must instead assemble rights from several, with the required set depending on the application and the country.¹,²
The fragmentation traces to an unresolved priority dispute over who first applied CRISPR-Cas9 to eukaryotic cells. The two most prominent groups are the University of California, Berkeley, the University of Vienna, and Emmanuelle Charpentier on one side, and the Broad Institute of MIT and Harvard on the other, with ToolGen and Sigma-Aldrich also holding foundational filings. The dispute has run through patent offices and courts for over a decade, and it remains live: in May 2025 the US Court of Appeals for the Federal Circuit vacated and remanded a decision that had awarded priority for eukaryotic CRISPR-Cas9 to the Broad Institute, reviving the Berkeley-led group's challenge.³,⁶ In Europe, the Berkeley-led group withdrew two foundational patents in late 2024 following an unfavorable preliminary opinion, then pursued divisional claims, while ToolGen secured European positions during 2025, illustrating how the landscape continues to shift among the competing groups.²,⁷ The academic literature has tracked this contested landscape since the technology's early years, documenting both its fragmentation and the licensing complexity it creates, and has examined proposed responses such as CRISPR patent pools.⁴,⁵,⁸
The practical consequence is that gene-editing FTO is a licensing-and-landscape problem, not a single clearance. The required rights differ by use, human therapeutics, agricultural and plant applications, research tools, and diagnostics can each implicate different foundational and improvement patents, and they differ by jurisdiction, because the same dispute has resolved differently in the United States, Europe, and Asia. The uncertainty is compounded by timing: some of the earliest, broadest patents may expire before the disputes are fully resolved, which shifts value toward the dense layer of improvement patents on delivery, specificity, and newer editing systems.² Because applications publish about eighteen months after filing, the most recent activity is under-represented, so the landscape is even larger and more active than granted-patent counts suggest.
Why CRISPR freedom-to-operate is hard
Fragmented foundational rights. Foundational claims are split across at least four groups, so clearing a product often requires multiple licenses rather than one.¹
Unresolved disputes. The priority dispute over eukaryotic CRISPR-Cas9 remains active, with a US Federal Circuit ruling in May 2025 reviving the Berkeley-led challenge, so ownership is not yet settled.³
Jurisdictional divergence. The same dispute has resolved differently across the United States, Europe, and Asia, so FTO must be assessed market-by-market.²
Application-specific rights. Human therapeutics, agriculture, research tools, and diagnostics implicate different patents, so the required license set depends on the intended use.¹
A dense improvement layer. Beyond the foundational patents, a large and growing layer of improvement patents covers delivery, specificity, base and prime editing, and newer nucleases, which is where much current FTO risk and white space now sit, including in application areas such as agricultural gene editing.⁹
How AI-powered landscape and FTO analysis helps
Navigating a landscape of more than twenty thousand families across multiple owners, applications, and jurisdictions is beyond manual search. AI-powered analysis addresses this with semantic search that retrieves relevant claims regardless of terminology, attribution that resolves owners to canonical entities so the fragmentation is visible, and continuous monitoring that tracks a fast-shifting landscape as disputes resolve and improvement patents publish. Because gene-editing advances appear in scientific literature before they are patented, reading both patents and literature gives earlier warning of where the improvement layer is extending.
Where Cypris fits
Cypris runs patent landscape and freedom-to-operate analysis for complex, fragmented fields such as gene editing across a corpus of more than 500 million patents and scientific papers, organized through a proprietary R&D ontology. The ontology clusters the landscape by technology and application and normalizes owners to canonical entities, so a team can see how foundational and improvement rights are distributed across the four groups and the many later filers rather than a flat list. Semantic search across patents and scientific literature surfaces relevant claims regardless of terminology and connects filings to the underlying research, which is where the improvement layer emerges first. Cypris Q, the platform's agentic layer, lets teams run landscape and FTO analysis conversationally and chain the attribution, clustering, and claim analysis, and Agentic Monitoring tracks the landscape over time and flags new filings and dispute developments as they publish. Cypris provides enterprise API partnerships with OpenAI, Anthropic, and Google, and is built with enterprise-grade security. Cypris serves hundreds of enterprise customers across pharmaceuticals, chemicals, advanced materials, energy, and other regulated industries.
FAQ
How large is the CRISPR patent landscape?
The CRISPR patent landscape is very large. A national patent office landscape analysis counted roughly 23,700 CRISPR patent families as of the end of 2024, up more than 6,500 in a single year. Because applications publish about eighteen months after filing, the most recent activity is under-represented, so the true landscape is even larger.
Why is freedom-to-operate hard for CRISPR?
Freedom-to-operate is hard for CRISPR because foundational rights are split across at least four competing groups and the key priority dispute remains unresolved, so a developer often cannot clear a product with a single license. The required rights also differ by application and jurisdiction. Assembling the correct set of licenses is the central FTO challenge.
What is the Broad versus UC Berkeley CRISPR dispute?
The Broad versus UC Berkeley dispute concerns who first applied CRISPR-Cas9 to eukaryotic cells, contested between the Berkeley-led group and the Broad Institute, with ToolGen and Sigma-Aldrich also holding foundational filings. In May 2025, the US Court of Appeals for the Federal Circuit vacated and remanded a decision that had favored the Broad Institute, reviving the Berkeley-led challenge. The dispute remains unresolved.
Does CRISPR freedom-to-operate differ by country?
Yes, CRISPR freedom-to-operate differs by country, because the same foundational dispute has resolved differently in the United States, Europe, and Asia. A party may hold stronger rights in one jurisdiction than another. FTO must therefore be assessed market-by-market rather than globally.
Why might a CRISPR product need multiple licenses?
A CRISPR product may need multiple licenses because foundational rights are fragmented across several owners, and improvement patents on delivery, specificity, and newer editing systems add further layers. The required set depends on the application and jurisdiction. This is why gene-editing FTO is a licensing-and-landscape problem rather than a single clearance.
How does the improvement-patent layer affect CRISPR FTO?
The improvement-patent layer affects CRISPR FTO because, beyond the foundational patents, a large and growing set of patents covers delivery, specificity, base and prime editing, and newer nucleases. As the earliest broad patents approach expiry, value shifts toward this layer, which is where much current FTO risk and white space sit. Mapping it requires reading both patents and scientific literature.
How does scientific literature help CRISPR landscape analysis?
Scientific literature helps CRISPR landscape analysis because gene-editing advances appear in research before they are patented, so the literature gives the earliest signal of where the improvement layer is extending. Analyzing patents alone gives a lagging view. Cypris analyzes both across more than 500 million patents and scientific papers.
Which teams need CRISPR patent landscape and FTO analysis?
CRISPR patent landscape and FTO analysis is needed by R&D, IP, and business-development teams in therapeutics, agriculture, industrial biotechnology, and diagnostics, along with investors assessing gene-editing assets. The fragmentation makes structured analysis essential. Cypris serves hundreds of enterprise customers across pharmaceuticals and other research-intensive industries.
How current does a CRISPR landscape need to be?
A CRISPR landscape needs to be continuously current, because the disputes are still resolving, new improvement patents publish constantly, and publication lag hides the most recent activity. A one-time landscape ages quickly. Cypris uses Agentic Monitoring to track the landscape and flag new filings and developments as they publish.
Endnotes
- Swiss Federal Institute of Intellectual Property (2025). CRISPR Technology: Patent & Licence Landscapes. https://www.ige.ch/
- Gowling WLG (2025). Fragmented and shifting CRISPR patent landscape: global proceedings and the patent pool solution. https://gowlingwlg.com/en/insights-resources/articles/2025/crispr-patent-landscape
- US Court of Appeals for the Federal Circuit (2025). Regents of the University of California v. Broad Institute, Inc., decided May 12, 2025. https://www.cafc.uscourts.gov/
- Egelie, K. J., Graff, G. D., Strand, S. P. & Johansen, B. (2016). The emerging patent landscape of CRISPR-Cas gene editing technology. Nature Biotechnology. https://doi.org/10.1038/nbt.3692
- Contreras, J. L. & Sherkow, J. S. (2017). CRISPR, surrogate licensing, and scientific discovery. Science. https://doi.org/10.1126/science.aal4222
- Sherkow, J. S. (2025). A "Bare Hope of a Result": The Second CRISPR Patent Appeal. The CRISPR Journal (analysis of the May 12, 2025 US Federal Circuit decision).
- Beck Greener (2025). Update on the CRISPR-Cas9 IP saga at the EPO: blows for both the Broad and CVC camps, but ToolGen ends 2025 with success. https://www.beckgreener.com/
- Stasi, A. & Pereira Rodrigues, I. (2019). Dealing with Patent Fragmentation in Genetics: Can Patent Pools Facilitate the Development of CRISPR Gene-Editing Technology? PubMed.
- Muhammad Adamu, U. et al. (2026). CRISPR in Wheat: Patents, Breeding Advances, and Emerging Challenges. Trends in Intellectual Property Research.

Agent orchestration in Microsoft Copilot works best when the orchestrator routes to scoped, governed connections rather than pulling every source into one undifferentiated context. The architecture that holds up under real R&D workloads keeps internal confidential data and external intelligence on separate trust boundaries, lets Copilot decide which to call, and treats external R&D and IP intelligence as a domain-oriented layer rather than a raw dataset dump. This guide explains how to design that orchestration so that a research team can ask a single question and have Copilot reason across an electronic lab notebook, internal developmental records, and the external patent and scientific literature without collapsing those very different data types into one fragile prompt.
Why orchestration belongs at the Copilot layer
The orchestrator is the component that decides which tool to call, in what order, and how to combine the results. In Microsoft Copilot Studio, generative orchestration is the mode that lets an agent select among multiple registered tools at runtime based on the user's intent and each tool's description. Microsoft requires generative orchestration to be enabled before an agent can use Model Context Protocol tools at all, which means the orchestration decision and the tool connections are designed to work as one system rather than as a hardcoded pipeline.
Putting orchestration at the Copilot layer matters for a specific reason. When orchestration is centralized, each connected source can stay narrow. The electronic lab notebook tool returns experimental records. The internal data tool returns developmental project context. The external intelligence tool returns patent and scientific findings. Copilot composes the answer from those scoped returns. The alternative, loading all of those corpora into a single context window and asking the model to sort it out, runs directly into context rot, the well-documented effect in which model accuracy degrades as the context window fills with more material. Centralized orchestration over scoped tools is the architectural answer to that degradation.
How MCP connections work inside Copilot Studio
Model Context Protocol is an open standard, introduced by Anthropic, that defines how applications expose tools and data to large language models in a consistent way. In Copilot Studio, MCP servers are made available through the same connector infrastructure that governs other Power Platform connections, which means an MCP connection inherits enterprise security and governance controls including Virtual Network integration, Data Loss Prevention policies, and multiple authentication methods.
Adding an MCP server to a Copilot Studio agent follows a defined path. From the agent's Tools page, you select Add a tool, then New tool, then Model Context Protocol, which opens the MCP onboarding wizard. You provide a server name, a server description, and a server URL, then select the authentication type the server requires. The server description is not cosmetic. The agent orchestrator reads that description at runtime to decide whether to call the server for a given user request, so a precise description of what each connection does is part of making orchestration work correctly. Once connected, each tool the MCP server publishes becomes an action inside Copilot Studio and inherits the server's defined inputs and outputs, and Copilot Studio reflects updates automatically as tools change on the server.
One governance fact shapes the entire design. Because MCP servers in Copilot Studio rely on Power Platform connectors for connectivity, any Data Loss Prevention policy that regulates those connectors also regulates the MCP server and its tools. This is the lever that lets a security team treat an internal ELN connection and an external intelligence connection under different policies even though both reach Copilot through the same mechanism.
Designing the internal trust boundary: ELN and developmental data
Internal confidential and developmental data is the most sensitive material in the orchestration, and it should be connected under the strictest governance. Electronic lab notebooks such as Benchling, LabArchives, and Scispot store the experimental records, sample data, and process documentation that represent a research organization's most valuable and proprietary information, and these platforms expose their data through documented REST APIs and emphasize regulatory compliance and data integrity as core features.
The design principle for this boundary is least exposure. The ELN connection and any internal developmental data connection should be governed by Data Loss Prevention policies that prevent confidential records from being combined with or transmitted to external destinations. Authentication should be scoped so the agent acts with the permissions of the requesting user rather than a broad service identity, which keeps the access model aligned with who is actually allowed to see which projects. Because Copilot Studio inherits connector-level DLP, a security team can place internal connections in a data group that is policy-isolated from external connections, so that the orchestrator can read from both but the platform enforces that confidential developmental data does not leak across the boundary. The internal tools should also be described narrowly to the orchestrator, so Copilot calls them only when a request genuinely concerns internal experimental or project data.
Designing the external boundary: patent and scientific intelligence
External R&D and IP intelligence is a fundamentally different kind of input, and treating it like just another data feed is where many agent designs go wrong. There is a meaningful difference between connecting an agent to a broad external dataset and connecting it to a domain-oriented intelligence layer. A raw external MCP endpoint that exposes a large patent or literature corpus hands the orchestrator an enormous, undifferentiated body of records, and asking the model to reason over that volume reintroduces the context rot problem the orchestration was meant to avoid. A domain-oriented layer instead returns a scoped, reasoned answer to the agent, so what enters Copilot's context is already a focused intelligence result rather than thousands of raw documents.
This is where the trust boundary and the quality boundary coincide. External intelligence should never share an undifferentiated context with confidential internal data, both because of data governance and because mixing a large external corpus into the same window as sensitive internal records degrades the reasoning on both. Keeping external intelligence as a separate, scoped connection that returns reasoned findings, rather than a firehose of raw records, protects accuracy and keeps the governance boundary clean.
Cypris as the external intelligence layer
This is the role Cypris is built for. As an enterprise R&D intelligence platform, Cypris unifies more than 500 million patents and scientific papers into a single intelligence layer with a proprietary R&D ontology, so that an agent reaching for external intelligence draws on the patent and scientific record in one reasoned place rather than across siloed connectors. Cypris is designed for R&D scientists and innovation strategists rather than IP attorneys, which means the intelligence it returns is scoped to the forward-looking questions research teams actually ask.
Crucially for an orchestration design, Cypris makes that intelligence available through official enterprise API partnerships with OpenAI, Anthropic, and Google, with enterprise-grade security built to Fortune 500 requirements. That partnership model lets the Cypris intelligence layer sit behind the AI tooling an organization already uses, including a Copilot orchestration, so the external intelligence entering the agent is a reasoned domain answer rather than a raw corpus. In the orchestration described here, Copilot routes external R&D and IP questions to Cypris as the domain-oriented intelligence layer, the internal ELN and developmental connections stay on their own governed boundary, and the orchestrator composes a single answer without ever collapsing confidential internal data and the external literature into one context. That separation is what makes the whole system both secure and accurate.
Putting the orchestration together
A working design has Copilot Studio as the orchestration layer with generative orchestration enabled, internal ELN and developmental data connected as narrowly scoped tools under isolating Data Loss Prevention policies, and external patent and scientific intelligence connected as a separate domain-oriented layer through Cypris's enterprise API partnerships. Each tool carries a precise description so the orchestrator routes correctly, authentication is scoped to the requesting user, and connector-level governance keeps the internal and external boundaries policy-separated. A researcher asks one question, and Copilot pulls scoped experimental context from the ELN, scoped project context from internal records, and a reasoned external intelligence answer from Cypris, then composes a response, all without ever forcing the model to reason over one bloated, mixed context. The result is an agent that is more accurate because each input is scoped and more secure because confidential developmental data never crosses into the external boundary.
FAQ
1. Can Microsoft Copilot orchestrate across both internal and external R&D data sources?Yes. Copilot Studio's generative orchestration mode lets a single agent select among multiple registered tools at runtime based on the user's intent, so one agent can route a question to an internal electronic lab notebook, internal developmental records, and an external intelligence layer and compose a unified answer.
2. What is generative orchestration in Copilot Studio?Generative orchestration is the mode in which the Copilot agent dynamically decides which tools to call and in what order based on the user's request and each tool's description, rather than following a hardcoded sequence. Microsoft requires it to be enabled before an agent can use Model Context Protocol tools.
3. How are MCP servers connected to a Copilot Studio agent?From the agent's Tools page you select Add a tool, then New tool, then Model Context Protocol, which opens the MCP onboarding wizard. You provide a server name, description, and URL, and select the authentication type. Each tool the server publishes becomes an action in Copilot Studio.
4. How is confidential R&D data kept secure in this architecture?MCP connections in Copilot Studio run on Power Platform connector infrastructure, so they inherit enterprise controls including Virtual Network integration, Data Loss Prevention policies, and multiple authentication methods. Internal connections can be placed under DLP policies that isolate them from external connections, and authentication can be scoped to the requesting user.
5. Why keep internal and external data on separate trust boundaries?Two reasons converge. Governance requires that confidential developmental data not leak to external destinations, and accuracy requires that a large external corpus not be mixed into the same context as sensitive internal records, because filling the context window with mixed material degrades the model's reasoning on both.
6. What is context rot and why does it matter for agent design?Context rot is the documented effect in which a model's accuracy declines as its context window fills with more material. It matters because loading multiple large corpora into one prompt, rather than routing to scoped tools, makes the agent reason worse, which is the core argument for centralizing orchestration over narrow connections.
7. How do electronic lab notebooks fit into the orchestration?ELN platforms such as Benchling, LabArchives, and Scispot hold experimental records, sample data, and process documentation, and expose that data through documented REST APIs. In the orchestration they are connected as narrowly scoped internal tools under strict governance, returning only the experimental context relevant to a given request.
8. What is the difference between connecting a raw external dataset and a domain-oriented intelligence layer?A raw external endpoint hands the orchestrator a large, undifferentiated body of records, which reintroduces context rot when the model tries to reason over the volume. A domain-oriented layer returns a scoped, reasoned answer, so what enters the agent's context is a focused result rather than thousands of raw documents.
9. How does Cypris connect into a Copilot orchestration?Cypris makes its R&D intelligence available through official enterprise API partnerships with OpenAI, Anthropic, and Google, with enterprise-grade security built to Fortune 500 requirements. That model lets the Cypris intelligence layer sit behind the AI tooling an organization already uses, so Copilot can route external patent and scientific questions to Cypris and receive a reasoned domain answer.
10. What does a complete orchestration design look like?Copilot Studio serves as the orchestration layer with generative orchestration enabled, internal ELN and developmental data are connected as scoped tools under isolating DLP policies, and external patent and scientific intelligence is connected as a separate domain-oriented layer through Cypris's enterprise API partnerships, with each tool precisely described so the orchestrator routes correctly.
Webinars

Many enterprises have adopted horizontal, foundation-model AI platforms. But access to the same underlying models does not, by itself, create differentiated intelligence. For highly technical and mission-critical research, general-purpose models may produce broad but weakly grounded answers when they lack access to authoritative technical data, specialized context, and verifiable sources.
The next competitive advantage will come from the intelligence layer surrounding the foundation model: the domain-specific data, ontologies, retrieval capabilities, agent workflows, and source grounding that together form an AI harness. These verticalized systems can transform general-purpose AI into a more specialized capability for research, innovation, and technical decision-making.
Join Steve Hafif, Co-Founder and CEO of Cypris.ai, and Marlene Valderrama, Principal IP Manager and Senior Technology Scout at Halliburton, for a conversation on the state of enterprise AI and how organizations can enhance horizontal AI platforms with verticalized intelligence designed for R&D and innovation.
.png)

Most IP organizations are making high-stakes capital allocation decisions with incomplete visibility – relying primarily on patent data as a proxy for innovation. That approach is not optimal. Patents alone cannot reveal technology trajectories, capital flows, or commercial viability.
A more effective model requires integrating patents with scientific literature, grant funding, market activity, and competitive intelligence. This means that for a complete picture, IP and R&D teams need infrastructure that connects fragmented data into a unified, decision-ready intelligence layer.
AI is accelerating that shift. The value is no longer simply in retrieving documents faster; it’s in extracting signal from noise. Modern AI systems can contextualize disparate datasets, identify patterns, and generate strategic narratives – transforming raw information into actionable insight.
Join us on Thursday, April 23, at 12 PM ET for a discussion on how unified AI platforms are redefining decision-making across IP and R&D teams. Moderated by Gene Quinn, panelists Marlene Valderrama and Amir Achourie will examine how integrating technical, scientific, and market data collapses traditional silos – enabling more aligned strategy, sharper investment decisions, and measurable business impact.
Register here: https://ipwatchdog.com/cypris-april-23-2026/
.png)
In this session, we break down how AI is reshaping the R&D lifecycle, from faster discovery to more informed decision-making. See how an intelligence layer approach enables teams to move beyond fragmented tools toward a unified, scalable system for innovation.
.avif)

%20-%20Next%20Generation%20Invisible%20Fishing%20Line.png)
%20-%20Noninvasive%20VNS.png)
%20-%20Low-Energy%20Desalination.png)