
Resources
Guides, research, and perspectives on R&D intelligence, IP strategy, and the future of AI enabled innovation.

Executive Summary
In 2024, US patent infringement jury verdicts totaled $4.19 billion across 72 cases. Twelve individual verdicts exceeded $100million. The largest single award—$857 million in General Access Solutions v.Cellco Partnership (Verizon)—exceeded the annual R&D budget of many mid-market technology companies. In the first half of 2025 alone, total damages reached an additional $1.91 billion.
The consequences of incomplete patent intelligence are not abstract. In what has become one of the most instructive IP disputes in recent history, Masimo’s pulse oximetry patents triggered a US import ban on certain Apple Watch models, forcing Apple to disable its blood oxygen feature across an entire product line, halt domestic sales of affected models, invest in a hardware redesign, and ultimately face a $634 million jury verdict in November 2025. Apple—a company with one of the most sophisticated intellectual property organizations on earth—spent years in litigation over technology it might have designed around during development.
For organizations with fewer resources than Apple, the risk calculus is starker. A mid-size materials company, a university spinout, or a defense contractor developing next-generation battery technology cannot absorb a nine-figure verdict or a multi-year injunction. For these organizations, the patent landscape analysis conducted during the development phase is the primary risk mitigation mechanism. The quality of that analysis is not a matter of convenience. It is a matter of survival.
And yet, a growing number of R&D and IP teams are conducting that analysis using general-purpose AI tools—ChatGPT, Claude, Microsoft Co-Pilot—that were never designed for patent intelligence and are structurally incapable of delivering it.
This report presents the findings of a controlled comparison study in which identical patent landscape queries were submitted to four AI-powered tools: Cypris (a purpose-built R&D intelligence platform),ChatGPT (OpenAI), Claude (Anthropic), and Microsoft Co-Pilot. Two technology domains were tested: solid-state lithium-sulfur battery electrolytes using garnet-type LLZO ceramic materials (freedom-to-operate analysis), and bio-based polyamide synthesis from castor oil derivatives (competitive intelligence).
The results reveal a significant and structurally persistent gap. In Test 1, Cypris identified over 40 active US patents and published applications with granular FTO risk assessments. Claude identified 12. ChatGPT identified 7, several with fabricated attribution. Co-Pilot identified 4. Among the patents surfaced exclusively by Cypris were filings rated as “Very High” FTO risk that directly claim the technology architecture described in the query. In Test 2, Cypris cited over 100 individual patent filings with full attribution to substantiate its competitive landscape rankings. No general-purpose model cited a single patent number.
The most active sectors for patent enforcement—semiconductors, AI, biopharma, and advanced materials—are the same sectors where R&D teams are most likely to adopt AI tools for intelligence workflows. The findings of this report have direct implications for any organization using general-purpose AI to inform patent strategy, competitive intelligence, or R&D investment decisions.

1. Methodology
A controlled comparative evaluation was conducted on March 27, 2026. An identical patent landscape query was submitted verbatim to each platform under standardized testing conditions. No follow-up prompts, clarifications, or iterative refinements were permitted, ensuring that each platform was evaluated based solely on its initial response.
The outputs were preserved in their original form and evaluated against predefined criteria using publicly verifiable patent records.
1.1 Query
Identify all active US patents and published applications filed in the last 5 years related to solid-state lithium-sulfur battery electrolytes using garnet-type ceramic materials. For each, provide the assignee, filing date, key claims, and current legal status. Highlight any patents that could pose freedom-to-operate risks for a company developing a Li₇La₃Zr₂O₁₂(LLZO)-based composite electrolyte with a polymer interlayer.
1.2 Tools Evaluated

1.3 Evaluation Criteria
Each response was evaluated using a consistent six-part scoring framework: patent coverage, assignee accuracy, filing metadata completeness, depth of claim analysis, quality of FTO risk stratification, and the presence of actionable strategic guidance.
Patent numbers, assignees, filing information, and legal status were independently checked against publicly available USPTO and WIPO records. The evaluation focused on the completeness, accuracy, and practical utility of each platform’s output rather than writing quality or presentation.
2. Findings
2.1 Coverage Gap
The most significant finding is the scale of the coverage differential. Cypris identified over 40 active US patents and published applications spanning LLZO-polymer composite electrolytes, garnet interface modification, polymer interlayer architectures, lithium-sulfur specific filings, and adjacent ceramic composite patents. The results were organized by technology category with per-patent FTO risk ratings.
Claude identified 12 patents organized in a four-tier risk framework. Its analysis was structurally sound and correctly flagged the two highest-risk filings (Solid Energies US 11,967,678 and the LLZO nanofiber multilayer US 11,923,501). It also identified the University ofMaryland/ Wachsman portfolio as a concentration risk and noted the NASA SABERS portfolio as a licensing opportunity. However, it missed the majority of the landscape, including the entire Corning portfolio, GM's interlayer patents, theKorea Institute of Energy Research three-layer architecture, and the HonHai/SolidEdge lithium-sulfur specific filing.
ChatGPT identified 7 patents, but the quality of attribution was inconsistent. It listed assignees as "Likely DOE /national lab ecosystem" and "Likely startup / defense contractor cluster" for two filings—language that indicates the model was inferring rather than retrieving assignee data. In a freedom-to-operate context, an unverified assignee attribution is functionally equivalent to no attribution, as it cannot support a licensing inquiry or risk assessment.
Co-Pilot identified 4 US patents. Its output was the most limited in scope, missing the Solid Energies portfolio entirely, theUMD/ Wachsman portfolio, Gelion/ Johnson Matthey, NASA SABERS, and all Li-S specific LLZO filings.
2.2 Critical Patents Missed by Public Models
The following table presents patents identified exclusively by Cypris that were rated as High or Very High FTO risk for the proposed technology architecture. None were surfaced by any general-purpose model.

2.3 Patent Fencing: The Solid Energies Portfolio
Cypris identified a coordinated patent fencing strategy by Solid Energies, Inc. that no general-purpose model detected at scale. Solid Energies holds at least four granted US patents and one published application covering LLZO-polymer composite electrolytes across compositions(US-12463245-B2), gradient architectures (US-12283655-B2), electrode integration (US-12463249-B2), and manufacturing processes (US-20230035720-A1). Claude identified one Solid Energies patent (US 11,967,678) and correctly rated it as the highest-priority FTO concern but did not surface the broader portfolio. ChatGPT and Co-Pilot identified zero Solid Energies filings.
The practical significance is that a company relying on any individual patent hit would underestimate the scope of Solid Energies' IP position. The fencing strategy—covering the composition, the architecture, the electrode integration, and the manufacturing method—means that identifying a single design-around for one patent does not resolve the FTO exposure from the portfolio as a whole. This is the kind of strategic insight that requires seeing the full picture, which no general-purpose model delivered
2.4 Assignee Attribution Quality
ChatGPT's response included at least two instances of fabricated or unverifiable assignee attributions. For US 11,367,895 B1, the listed assignee was "Likely startup / defense contractor cluster." For US 2021/0202983 A1, the assignee was described as "Likely DOE / national lab ecosystem." In both cases, the model appears to have inferred the assignee from contextual patterns in its training data rather than retrieving the information from patent records.
In any operational IP workflow, assignee identity is foundational. It determines licensing strategy, litigation risk, and competitive positioning. A fabricated assignee is more dangerous than a missing one because it creates an illusion of completeness that discourages further investigation. An R&D team receiving this output might reasonably conclude that the landscape analysis is finished when it is not.
3. Structural Limitations of General-Purpose Models for Patent Intelligence
3.1 Training Data Is Not Patent Data
Large language models are trained on web-scraped text. Their knowledge of the patent record is derived from whatever fragments appeared in their training corpus: blog posts mentioning filings, news articles about litigation, snippets of Google Patents pages that were crawlable at the time of data collection. They do not have systematic, structured access to the USPTO database. They cannot query patent classification codes, parse claim language against a specific technology architecture, or verify whether a patent has been assigned, abandoned, or subjected to terminal disclaimer since their training data was collected.
This is not a limitation that improves with scale. A larger training corpus does not produce systematic patent coverage; it produces a larger but still arbitrary sampling of the patent record. The result is that general-purpose models will consistently surface well-known patents from heavily discussed assignees (QuantumScape, for example, appeared in most responses) while missing commercially significant filings from less publicly visible entities (Solid Energies, Korea Institute of EnergyResearch, Shenzhen Solid Advanced Materials).
3.2 The Web Is Closing to Model Scrapers
The data access problem is structural and worsening. As of mid-2025, Cloudflare reported that among the top 10,000 web domains, the majority now fully disallow AI crawlers such as GPTBot andClaudeBot via robots.txt. The trend has accelerated from partial restrictions to outright blocks, and the crawl-to-referral ratios reveal the underlying tension: OpenAI's crawlers access approximately1,700 pages for every referral they return to publishers; Anthropic's ratio exceeds 73,000 to 1.
Patent databases, scientific publishers, and IP analytics platforms are among the most restrictive content categories. A Duke University study in 2025 found that several categories of AI-related crawlers never request robots.txt files at all. The practical consequence is that the knowledge gap between what a general-purpose model "knows" about the patent landscape and what actually exists in the patent record is widening with each training cycle. A landscape query that a general-purpose model partially answered in 2023 may return less useful information in 2026.
3.3 General-Purpose Models Lack Ontological Frameworks for Patent Analysis
A freedom-to-operate analysis is not a summarization task. It requires understanding claim scope, prosecution history, continuation and divisional chains, assignee normalization (a single company may appear under multiple entity names across patent records), priority dates versus filing dates versus publication dates, and the relationship between dependent and independent claims. It requires mapping the specific technical features of a proposed product against independent claim language—not keyword matching.
General-purpose models do not have these frameworks. They pattern-match against training data and produce outputs that adopt the format and tone of patent analysis without the underlying data infrastructure. The format is correct. The confidence is high. The coverage is incomplete in ways that are not visible to the user.
4. Comparative Output Quality
The following table summarizes the qualitative characteristics of each tool's response across the dimensions most relevant to an operational IP workflow.

5. Implications for R&D and IP Organizations
5.1 The Confidence Problem
The central risk identified by this study is not that general-purpose models produce bad outputs—it is that they produce incomplete outputs with high confidence. Each model delivered its results in a professional format with structured analysis, risk ratings, and strategic recommendations. At no point did any model indicate the boundaries of its knowledge or flag that its results represented a fraction of the available patent record. A practitioner receiving one of these outputs would have no signal that the analysis was incomplete unless they independently validated it against a comprehensive datasource.
This creates an asymmetric risk profile: the better the format and tone of the output, the less likely the user is to question its completeness. In a corporate environment where AI outputs are increasingly treated as first-pass analysis, this dynamic incentivizes under-investigation at precisely the moment when thoroughness is most critical.
5.2 The Diversification Illusion
It might be assumed that running the same query through multiple general-purpose models provides validation through diversity of sources. This study suggests otherwise. While the four tools returned different subsets of patents, all operated under the same structural constraints: training data rather than live patent databases, web-scraped content rather than structured IP records, and general-purpose reasoning rather than patent-specific ontological frameworks. Running the same query through three constrained tools does not produce triangulation; it produces three partial views of the same incomplete picture.
5.3 The Appropriate Use Boundary
General-purpose language models are effective tools for a wide range of tasks: drafting communications, summarizing documents, generating code, and exploratory research. The finding of this study is not that these tools lack value but that their value boundary does not extend to decisions that carry existential commercial risk.
Patent landscape analysis, freedom-to-operate assessment, and competitive intelligence that informs R&D investment decisions fall outside that boundary. These are workflows where the completeness and verifiability of the underlying data are not merely desirable but are the primary determinant of whether the analysis has value. A patent landscape that captures 10% of the relevant filings, regardless of how well-formatted or confidently presented, is a liability rather than an asset.
6. Test 2: Competitive Intelligence — Bio-Based Polyamide Patent Landscape
To assess whether the findings from Test 1 were specific to a single technology domain or reflected a broader structural pattern, a second query was submitted to all four tools. This query shifted from freedom-to-operate analysis to competitive intelligence, asking each tool to identify the top 10organizations by patent filing volume in bio-based polyamide synthesis from castor oil derivatives over the past three years, with summaries of technical approach, co-assignee relationships, and portfolio trajectory.
6.1 Query

6.2 Summary of Results

6.3 Key Differentiators
Verifiability
The most consequential difference in Test 2 was the presence or absence of verifiable evidence. Cypris cited over 100 individual patent filings with full patent numbers, assignee names, and publication dates. Every claim about an organization’s technical focus, co-assignee relationships, and filing trajectory was anchored to specific documents that a practitioner could independently verify in USPTO, Espacenet, or WIPO PATENT SCOPE. No general-purpose model cited a single patent number. Claude produced the most structured and analytically useful output among the public models, with estimated filing ranges, product names, and strategic observations that were directionally plausible. However, without underlying patent citations, every claim in the response requires independent verification before it can inform a business decision. ChatGPT and Co-Pilot offered thinner profiles with no filing counts and no patent-level specificity.
Data Integrity
ChatGPT’s response contained a structural error that would mislead a practitioner: it listed CathayBiotech as organization #5 and then listed “Cathay Affiliate Cluster” as a separate organization at #9, effectively double-counting a single entity. It repeated this pattern with Toray at #4 and “Toray(Additional Programs)” at #10. In a competitive intelligence context where the ranking itself is the deliverable, this kind of error distorts the landscape and could lead to misallocation of competitive monitoring resources.
Organizations Missed
Cypris identified Kingfa Sci. & Tech. (8–10 filings with a differentiated furan diacid-based polyamide platform) and Zhejiang NHU (4–6 filings focused on continuous polymerization process technology)as emerging players that no general-purpose model surfaced. Both represent potential competitive threats or partnership opportunities that would be invisible to a team relying on public AI tools.Conversely, ChatGPT included organizations such as ANTA and Jiangsu Taiji that appear to be downstream users rather than significant patent filers in synthesis, suggesting the model was conflating commercial activity with IP activity.
Strategic Depth
Cypris’s cross-cutting observations identified a fundamental chemistry divergence in the landscape:European incumbents (Arkema, Evonik, EMS) rely on traditional castor oil pyrolysis to 11-aminoundecanoic acid or sebacic acid, while Chinese entrants (Cathay Biotech, Kingfa) are developing alternative bio-based routes through fermentation and furandicarboxylic acid chemistry.This represents a potential long-term disruption to the castor oil supply chain dependency thatWestern players have built their IP strategies around. Claude identified a similar theme at a higher level of abstraction. Neither ChatGPT nor Co-Pilot noted the divergence.
6.4 Test 2 Conclusion
Test 2 confirms that the coverage and verifiability gaps observed in Test 1 are not domain-specific.In a competitive intelligence context—where the deliverable is a ranked landscape of organizationalIP activity—the same structural limitations apply. General-purpose models can produce plausible-looking top-10 lists with reasonable organizational names, but they cannot anchor those lists to verifiable patent data, they cannot provide precise filing volumes, and they cannot identify emerging players whose patent activity is visible in structured databases but absent from the web-scraped content that general-purpose models rely on.
7. Conclusion
This comparative analysis, spanning two distinct technology domains and two distinct analytical workflows—freedom-to-operate assessment and competitive intelligence—demonstrates that the gap between purpose-built R&D intelligence platforms and general-purpose language models is not marginal, not domain-specific, and not transient. It is structural and consequential.
In Test 1 (LLZO garnet electrolytes for Li-S batteries), the purpose-built platform identified more than three times as many patents as the best-performing general-purpose model and ten times as many as the lowest-performing one. Among the patents identified exclusively by the purpose-built platform were filings rated as Very High FTO risk that directly claim the proposed technology architecture. InTest 2 (bio-based polyamide competitive landscape), the purpose-built platform cited over 100individual patent filings to substantiate its organizational rankings; no general-purpose model cited as ingle patent number.
The structural drivers of this gap—reliance on training data rather than live patent feeds, the accelerating closure of web content to AI scrapers, and the absence of patent-specific analytical frameworks—are not transient. They are inherent to the architecture of general-purpose models and will persist regardless of increases in model capability or training data volume.
For R&D and IP leaders, the practical implication is clear: general-purpose AI tools should be used for general-purpose tasks. Patent intelligence, competitive landscaping, and freedom-to-operate analysis require purpose-built systems with direct access to structured patent data, domain-specific analytical frameworks, and the ability to surface what a general-purpose model cannot—not because it chooses not to, but because it structurally cannot access the data.
The question for every organization making R&D investment decisions today is whether the tools informing those decisions have access to the evidence base those decisions require. This study suggests that for the majority of general-purpose AI tools currently in use, the answer is no.
Study Disclosure
This comparative evaluation was commissioned and published by Cypris. The testing methodology, prompts, evaluation criteria, and underlying outputs have been documented to support independent review and replication.
All platform outputs were preserved in their original form. Patent data and material factual claims were cross-checked against USPTO Patent Center and WIPO PATENTSCOPE records as of March 27, 2026. Cypris was one of the platforms evaluated and therefore has a commercial interest in the findings.
The Patent Intelligence Gap - A Comparative Analysis of Verticalized AI-Patent Tools vs. General-Purpose Language Models for R&D Decision-Making
Blogs

Prior art search for artificial intelligence and machine learning inventions is one of the hardest retrieval problems in patent work, for reasons specific to how AI knowledge is produced and disclosed. Prior art search establishes whether an invention is novel by finding any earlier disclosure that describes it. In most fields, the relevant disclosures are predominantly patents. In AI and machine learning, the most relevant and most recent disclosures are predominantly non-patent literature: preprints on arXiv, proceedings from conferences such as NeurIPS and ICML, open-source code and model documentation, and technical reports. These sources are published quickly and openly, often well ahead of any corresponding patent, so a prior art search confined to patent databases misses the state of the art.
The volume compounds the difficulty. AI scientific publications more than doubled from about 102,000 in 2013 to more than 242,000 in 2023, growing nearly 20 percent in the final year alone.¹ Patenting has grown even faster from a smaller base: AI patents granted worldwide rose from 3,833 in 2010 to 122,511 in 2023, an increase of almost 30 percent in the last year measured.¹ Generative AI illustrates the velocity of the literature most sharply, with related scientific publications rising from 116 in 2014 to more than 34,000 in 2023 while generative-AI patent families grew more than 800 percent over roughly the same period.² A prior art searcher in this field is therefore working against both a large and a rapidly expanding corpus, split across patent and non-patent sources.
Retrieval quality falls exactly where AI prior art needs it most. Patent retrieval is already harder than general-domain information retrieval, and controlled evaluation shows that cross-domain retrieval, finding relevant art outside the query's own technology area, performs several times worse than in-domain retrieval; one recent family-level benchmark found out-of-domain retrieval roughly five times worse than in-domain across hundreds of controlled configurations.³,⁴ AI and machine-learning methods are applied across many application domains, so relevant prior art for an AI invention is frequently located in a different field than the invention's stated use, which is precisely the cross-domain case where conventional retrieval degrades. This is the technical reason keyword and classification search alone are insufficient for AI prior art, and why dense, semantic methods have become the focus of research on patent prior art retrieval.⁵,⁶
Why AI prior art is distinctively hard
Non-patent literature dominates. The most relevant and most recent AI disclosures appear first in preprints, conference proceedings, and open-source code, so a patent-only search misses the state of the art.
Exploding volume. AI publications more than doubled to over 242,000 in 2023, and AI patents granted rose to 122,511, so the corpus a searcher must cover is both large and expanding rapidly.¹
Cross-domain dispersion. AI methods are applied across many fields, so relevant prior art is often in a different technology area than the invention, which is where retrieval degrades most.³
Fast obsolescence of terminology. AI vocabulary evolves quickly, so keyword search misses conceptually identical work described in newer or different terms.
Software-claim breadth. Algorithmic and software claims can be drafted broadly and abstractly, which makes matching a claim to its closest prior art a conceptual rather than a lexical task.
How semantic search closes the gap
Semantic search addresses each of these problems. It retrieves conceptually relevant disclosures regardless of terminology, which handles both fast-evolving vocabulary and broadly drafted software claims. Applied across both patents and scientific literature in one corpus, it covers the non-patent literature where AI prior art concentrates rather than patents alone. And because dense retrieval encodes meaning rather than surface form, it is better positioned than keyword search for the cross-domain case, retrieving relevant art from a different application area than the invention. Combined with an ontology that organizes retrieval by concept, semantic search returns a structured, high-recall view of the prior art rather than a keyword-limited sample.
Where Cypris fits
Cypris runs semantic prior art search across a corpus of more than 500 million patents and scientific papers, organized through a proprietary R&D ontology. Because the corpus spans both patents and scientific literature, Cypris covers the non-patent literature where AI and machine-learning prior art concentrates, rather than patents alone. Semantic search retrieves conceptually relevant disclosures regardless of terminology, which handles the fast-evolving vocabulary and broadly drafted software claims characteristic of AI inventions, and the ontology organizes retrieval by concept so cross-domain prior art in a different application area is surfaced rather than missed. Cypris Q, the platform's agentic layer, lets teams run and chain prior art and novelty analysis conversationally, and Agentic Monitoring tracks a technology area over time so newly published disclosures are surfaced as they appear, which matters in a field moving as fast as AI. Cypris provides enterprise API partnerships with OpenAI, Anthropic, and Google, and is built with enterprise-grade security. Cypris serves hundreds of enterprise customers across pharmaceuticals, chemicals, advanced materials, energy, and other regulated industries.
FAQ
Why is prior art search hard for AI and machine learning inventions?
Prior art search is hard for AI and machine learning inventions because the most relevant and most recent disclosures are predominantly non-patent literature, such as preprints, conference proceedings, and open-source code, which a patent-only search misses. The corpus is also large and expanding rapidly, and AI methods are dispersed across many application domains. These factors make high-recall, cross-domain retrieval essential.
Why does non-patent literature matter so much for AI prior art?
Non-patent literature matters for AI prior art because AI research is published quickly and openly, often well ahead of any corresponding patent, so the state of the art appears first in preprints, conference papers, and code. A search confined to patent databases misses these disclosures. Effective AI prior art search must cover both patents and scientific literature.
How large is the AI prior art corpus?
The AI prior art corpus is large and growing quickly. AI scientific publications more than doubled from about 102,000 in 2013 to over 242,000 in 2023, and AI patents granted worldwide rose from 3,833 in 2010 to 122,511 in 2023. Generative-AI publications alone grew from 116 in 2014 to more than 34,000 in 2023.
What makes AI prior art retrieval technically difficult?
AI prior art retrieval is technically difficult because AI methods are applied across many domains, so relevant prior art is often in a different technology area than the invention, and cross-domain retrieval performs several times worse than in-domain retrieval. One benchmark found out-of-domain retrieval roughly five times worse than in-domain. Fast-evolving terminology and broadly drafted software claims add further difficulty.
Why is keyword search insufficient for AI prior art?
Keyword search is insufficient for AI prior art because AI terminology evolves quickly and software claims are often drafted broadly and abstractly, so conceptually identical work is described in different terms. Keyword search matches surface form and misses these. Semantic search retrieves by meaning, which is what the task requires.
How does semantic search improve AI prior art search?
Semantic search improves AI prior art search by retrieving conceptually relevant disclosures regardless of terminology, across both patents and scientific literature, and by handling the cross-domain case where relevant art is in a different field. It encodes meaning rather than surface form. Combined with an ontology, it returns a structured, high-recall view of the prior art.
Does AI prior art search need to cover scientific literature?
AI prior art search needs to cover scientific literature because the most relevant and most recent AI disclosures appear there first, in preprints, conference proceedings, and technical reports. Covering patents alone leaves the state of the art unretrieved. Cypris searches both across more than 500 million patents and scientific papers.
Which teams run AI prior art search?
AI prior art search is run by IP, R&D, and patent teams at technology companies and across industries adopting AI, as well as by patent professionals assessing novelty. It is increasingly important as AI patenting grows. Cypris serves hundreds of enterprise customers across research-intensive and regulated industries.
How current does AI prior art search need to be?
AI prior art search needs to be continuously current, because AI research and filings publish constantly and the state of the art shifts quickly. A one-time search reflects only the moment it was run. Cypris uses Agentic Monitoring to track a technology area and surface newly published disclosures as they appear.
Endnotes
- Stanford Institute for Human-Centered Artificial Intelligence (2025). Artificial Intelligence Index Report 2025, Chapter 1. arXiv:2504.07139. https://doi.org/10.48550/arxiv.2504.07139
- World Intellectual Property Organization (2024). Patent Landscape Report: Generative Artificial Intelligence. Geneva: WIPO. https://doi.org/10.34667/tind.49740
- Cavallucci, N., Chibane, I. & Ayaou, M. (2026). DAPFAM: A Domain-Aware Family-level Dataset to benchmark cross-domain patent retrieval. Array. https://doi.org/10.1016/j.array.2026.100720
- Lupu, M. (2013). Patent Retrieval. Foundations and Trends in Information Retrieval. https://doi.org/10.1561/1500000027
- Stamatis, V. (2022). End to End Neural Retrieval for Patent Prior Art Search. Lecture Notes in Computer Science. https://doi.org/10.1007/978-3-030-99739-7_66
- Zihayat, M. & Etwaroo, R. (2021). A non-factoid question answering system for prior art search. Expert Systems with Applications. https://doi.org/10.1016/j.eswa.2021.114910

Tightening regulation of per- and polyfluoroalkyl substances is reshaping materials chemistry, and it is opening patent white space for organizations that can develop fluorine-free alternatives. PFAS are used for water, oil, and stain resistance across coatings, textiles, firefighting foams, membranes, semiconductors, and food packaging, and they are now the subject of the broadest chemical restriction ever proposed in Europe. The universal PFAS restriction proposal submitted to the European Chemicals Agency in January 2023 by five national authorities covers on the order of 10,000 substances, and it drew more than 5,600 comments from over 4,400 organizations, an unprecedented response that reflects how many industries are affected.¹ The scope depends on definition: under the 2021 OECD definition, which classifies a substance as PFAS if it contains at least one fully fluorinated carbon, several million catalogued substances qualify, while the number in active commercial use is far smaller.²
The regulatory trajectory is a sequence of tightening actions rather than a single event, which is what makes the resulting innovation demand durable. In the European Union, restrictions moved from PFOS in 2006 to PFOA and related substances in later years, to a PFHxA restriction adopted in 2024, a ban on PFAS in firefighting foams, and a ban on PFAS in food-contact packaging taking effect in 2026, with the universal restriction proposal under scientific evaluation through 2026.¹ In the United States, the Environmental Protection Agency finalized the first national drinking-water limits for several PFAS in 2024, setting maximum contaminant levels of 4.0 parts per trillion for PFOA and PFOS and higher limits for other compounds, and designated PFOA and PFOS as hazardous substances under the federal cleanup statute the same year.³ ECHA has estimated that, absent action, several million tonnes of PFAS would reach the environment over the coming decades.¹
This regulatory pressure is a well-understood driver of innovation. The Porter hypothesis, that well-designed environmental regulation can induce innovation that partly or wholly offsets compliance costs, has been supported across two decades of evidence and a multi-country meta-analysis, and firm-level studies show environmental regulation inducing greener product innovation specifically in chemical industries.⁴,⁵,⁶ For materials developers, the implication is direct: regulation is converting fluorine-free chemistry from a niche into a competitive frontier, and the organizations that build defensible IP positions early will hold advantage as substitution accelerates.
Where the white space is
Firefighting foams. Fluorine-free foams are the most advanced substitution area, driven by bans on PFAS-containing aqueous film-forming foams, though performance and toxicity gaps relative to legacy foams remain an active research and patenting frontier.⁷
Textile and coating treatments. Water- and oil-repellent finishes are a major PFAS use, and fluorine-free hydrophobic and oleophobic coatings, including bio-based and hierarchical-structured approaches, are an active area of development with room for defensible positions.⁸,⁹
Membranes and packaging. Food-contact packaging faces near-term bans, and membrane and barrier applications require substitutes that match performance, which keeps white space open where a fluorine-free chemistry can meet the functional requirement.
Semiconductors and specialty uses. Certain high-performance uses have few current substitutes, so these areas are simultaneously the hardest to displace and the most valuable to solve, and the patent landscape around viable alternatives is comparatively sparse.
How to find PFAS-alternative white space
Scope the application area and functional requirement precisely, since PFAS substitution is application-specific and a fluorine-free chemistry that works for textiles may not work for firefighting foam.
Map patents and scientific literature across the fluorine-free chemistries relevant to that application, because materials research precedes patenting and gives the earliest signal of a viable alternative.
Cluster activity by concept and attribute it to organizations, using an ontology to group related chemistry and normalize assignees, so dense and sparse regions are visible.
Identify the sparse, defensible regions, distinguishing genuine white space from areas that are sparse only because a chemistry does not yet meet the functional requirement.
Monitor continuously, tracking both the chemistry and the regulatory timeline, so filings and restrictions are surfaced as they publish and a white space position is secured before substitution accelerates.
Where Cypris fits
Cypris runs patent landscape and white space analysis for regulation-driven fields such as PFAS alternatives across a corpus of more than 500 million patents and scientific papers, organized through a proprietary R&D ontology. Semantic search across patents and scientific literature surfaces fluorine-free chemistry regardless of nomenclature and connects filings to the underlying materials research, which is where the earliest signals of viable alternatives appear. The ontology clusters activity by application and chemistry and normalizes organizations to canonical entities, so a team can resolve which fluorine-free approaches are crowded and which remain open as white space. Cypris Q, the platform's agentic layer, lets teams run landscape and white space analysis conversationally and chain the search, attribution, and gap analysis, and Agentic Monitoring tracks a defined chemistry over time and flags new filings as they publish. Cypris provides enterprise API partnerships with OpenAI, Anthropic, and Google, and is built with enterprise-grade security. Cypris serves hundreds of enterprise customers across pharmaceuticals, chemicals, advanced materials, energy, and other regulated industries.
FAQ
Why is PFAS regulation creating patent white space?
PFAS regulation is creating patent white space by driving demand for fluorine-free alternatives across many applications, faster than defensible IP positions have been established. The EU REACH universal restriction covers roughly 10,000 substances and US EPA rules now limit several PFAS, so substitution is accelerating. Organizations that build fluorine-free IP early can hold advantage as demand rises.
What is the EU REACH universal PFAS restriction?
The EU REACH universal PFAS restriction is a proposal submitted to the European Chemicals Agency in January 2023 by five national authorities to restrict the manufacture and use of PFAS as a class, covering on the order of 10,000 substances. It drew more than 5,600 comments from over 4,400 organizations. It is under scientific evaluation, with the outcome expected to shape substitution across many industries.
What US rules apply to PFAS?
In the United States, the Environmental Protection Agency finalized the first national drinking-water limits for several PFAS in 2024, setting maximum contaminant levels of 4.0 parts per trillion for PFOA and PFOS and higher limits for other compounds, and designated PFOA and PFOS as hazardous substances under the federal cleanup statute the same year. These actions increase the pressure to substitute PFAS. They apply alongside state-level restrictions.
Which application areas have the most PFAS-alternative white space?
Firefighting foams, textile and coating treatments, membranes and packaging, and certain semiconductor and specialty uses all have PFAS-alternative white space, though the amount varies. Firefighting-foam alternatives are the most advanced, while high-performance specialty uses have few substitutes and are the most valuable to solve. White space is largest where a fluorine-free chemistry can meet the functional requirement but few patents yet exist.
How does regulation drive innovation in materials?
Regulation drives innovation in materials by creating demand for compliant substitutes, a pattern described by the Porter hypothesis and supported by two decades of evidence and firm-level studies in chemical industries. Well-designed regulation induces innovation that can partly offset compliance costs. For PFAS, this is converting fluorine-free chemistry from a niche into a competitive frontier.
How do you find white space in PFAS alternatives?
Finding white space in PFAS alternatives means scoping a specific application and functional requirement, mapping patents and scientific literature across the relevant fluorine-free chemistries, clustering activity by concept, and identifying the sparse, defensible regions. Because materials research precedes patenting, literature coverage gives early signal. The analysis must distinguish genuine white space from areas that are sparse because no chemistry yet meets the requirement.
Why does PFAS-alternative analysis need scientific literature?
PFAS-alternative analysis needs scientific literature because fluorine-free chemistries appear in research before they are patented, so the literature gives the earliest signal of a viable alternative. Analyzing patents alone gives a lagging view. Cypris analyzes both across more than 500 million patents and scientific papers.
Which teams work on PFAS alternatives?
PFAS alternatives are developed by R&D, innovation, and IP teams in chemicals, advanced materials, coatings, textiles, consumer products, and their suppliers, alongside regulatory affairs. The work is driven by tightening regulation and customer demand for fluorine-free products. Cypris serves hundreds of enterprise customers across chemicals, advanced materials, and other regulated industries.
How do you keep a PFAS-alternatives landscape current? Keeping a PFAS-alternatives landscape current requires continuous monitoring of both the chemistry and the regulatory timeline, because filings and restrictions evolve constantly. A one-time landscape ages quickly as new rules and patents publish. Cypris uses Agentic Monitoring to track a defined chemistry over time and flag new filings as they publish.
Endnotes
- European Chemicals Agency. Registry of restriction intentions: per- and polyfluoroalkyl substances (PFAS) universal restriction proposal (2023) and related consultation and evaluation materials. https://echa.europa.eu/
- OECD (2021). Reconciling Terminology of the Universe of Per- and Polyfluoroalkyl Substances: Recommendations and Practical Guidance; and US Environmental Protection Agency PFAS inventory materials.
- US Environmental Protection Agency (2024). PFAS National Primary Drinking Water Regulation; and CERCLA designation of PFOA and PFOS as hazardous substances. https://www.epa.gov/pfas
- Ambec, S., Cohen, M. A., Elgie, S. & Lanoie, P. (2013). The Porter Hypothesis at 20. Review of Environmental Economics and Policy. https://doi.org/10.1093/reep/res016
- Yan, Z., Li, Y., Zhang, X. & Zhu, J. (2024). Revisiting the Porter hypothesis: a multi-country meta-analysis. Humanities and Social Sciences Communications. https://doi.org/10.1057/s41599-024-02671-9
- Choi, J., Kang, J. & Chung, S. (2025). Environmental regulation, induced innovation, and greener transition: firm-level evidence. Journal of Development Economics. https://doi.org/10.1016/j.jdeveco.2025.103678
- Hossain, T., Ormond, R. B. et al. (2024). Exploring the Prospects and Challenges of Fluorine-Free Firefighting Foams (F3) as Alternatives to AFFF: A Review. ACS Omega. https://doi.org/10.1021/acsomega.4c03673
- Likozar, B. et al. (2024). Unveiling PFAS-free Solutions for Hydrophobic and Oleophobic Textile Coatings. https://doi.org/10.55295/psl.2024.i19
- Nicolas, M. et al. (2024). PFAS-free hierarchical superhydrophobic textiles. Advanced Engineering Materials. https://doi.org/10.1002/adem.202401736

Freedom-to-operate for GLP-1 receptor agonists and peptide therapeutics is among the most demanding FTO problems in pharmaceuticals, because protection in this class is built as a dense, layered thicket that extends far beyond the active ingredient. Freedom-to-operate determines whether making, using, or selling a product would infringe another party's active patent claims. In the GLP-1 and peptide space, answering that question requires reading many claim types across many patents, because a single product is protected by a stack of filings covering the molecule, its formulation, its dosing, its delivery device, and its manufacture. A peer-reviewed analysis of GLP-1 receptor agonists approved between 2005 and 2021 found that manufacturers listed a median of 19.5 patents per product, that 54 percent of those patents were on delivery devices rather than the active ingredient, that the median expected protection was 18.3 years after approval, and that no generic manufacturer had yet successfully challenged a GLP-1 receptor agonist patent.¹
The commercial stakes are large. Industry analyst forecasts vary widely with scope, placing the GLP-1 market anywhere from the low tens of billions of dollars to well over one hundred billion by 2030 and projecting double-digit annual growth; these are analyst estimates rather than authoritative figures, and they differ mainly in what they count.² The scale of the opportunity is what drives the density of the patent thicket, because each additional protected feature can delay competition on a high-revenue product. For any organization developing a follow-on peptide, a biosimilar, or a differentiated GLP-1 product, FTO is therefore a gating analysis rather than a formality.
Peptide therapeutics compound the difficulty. Peptides can be claimed as sequences and modifications, formulated for stability and half-life extension, delivered by injection or increasingly by oral routes, and manufactured through distinct synthesis and purification processes, so the claim surface is broad. Recent filing activity has shifted toward oral delivery, dual and triple receptor agonists, and combination therapies, which is where both the newest FTO risk and the remaining white space now sit.³ An FTO analysis in this class has to cover all of these dimensions, and it has to stay current as the frontier moves.
What creates FTO risk in GLP-1 and peptide products
Composition-of-matter claims. These cover the peptide itself, including sequences, analogues, and modifications, and are the primary protection, though in a mature class many core molecules approach expiry.
Formulation claims. These cover stabilized, extended-release, and oral formulations, which are heavily patented, as formulation is where much peptide innovation and differentiation occurs.
Dosing-regimen and method-of-use claims. These cover titration schedules and specific therapeutic uses, and can block a product for a particular indication or regimen even when the molecule is otherwise available.
Delivery-device claims. These cover injection pens and other devices and are a large share of the thicket; peer-reviewed analysis found delivery devices accounted for the majority of listed GLP-1 patents and function as a distinct barrier to entry.¹,⁴
Process and manufacturing claims. These cover synthesis and purification routes, so a developer can be free to use a molecule yet blocked from a particular manufacturing method.
A single molecule illustrates the layering. A published patent landscape of one dual GLP-1/glucagon receptor agonist identified twelve patent families spanning composition-of-matter, process chemistry, formulation, dosing regimen, and method-of-use, a clean worked example of how all five claim types stack on one product.⁵
The thicket dynamic and the expiry landscape
The density of GLP-1 protection reflects a broader pharmaceutical pattern. Empirical analysis shows the number of patents filed per active ingredient rose from 1.86 in 2001 to nearly six by 2019, driven substantially by continuation applications, which account for roughly a third of small-molecule pharmaceutical patents.⁶ These secondary filings extend the effective protection period, and the economics of that extension, including how patent challenges and settlements shape effective market life, are well documented.⁷ Pharmaceutical thickets also differ structurally from thickets in complex-technology industries, which is why FTO methods developed for electronics do not transfer cleanly to peptides.⁸
The expiry landscape is the other half of the picture. As core molecules approach the end of composition-of-matter protection, the surrounding formulation, device, and process claims determine when and where competition can actually enter. Analysis of one leading GLP-1 molecule found that the timing of primary-patent expiry varies substantially by market, so freedom-to-operate for a follow-on product is jurisdiction-specific, and the practical entry date is governed by the secondary thicket rather than the headline molecule expiry.⁹ For a developer, this means FTO must be assessed claim-by-claim and market-by-market, not at the level of the molecule.
How AI-powered FTO helps
Navigating a thicket of this density by manual search is slow and prone to coverage gaps, which are the main source of FTO risk. AI-powered FTO addresses this with semantic search that retrieves relevant claims regardless of terminology, claim-level analysis that focuses on the independent claims defining infringement scope across all five claim types, and continuous monitoring that keeps a cleared position current as new formulation, device, and combination filings publish. Because peptide innovation appears in scientific literature before it is patented, reading both patents and literature gives earlier warning of where the thicket is extending.
Where Cypris fits
Cypris runs claim-level, semantic, AI-powered freedom-to-operate across a corpus of more than 500 million patents and scientific papers, organized through a proprietary R&D ontology. Semantic search across patents and scientific literature surfaces relevant claims regardless of terminology, across composition, formulation, dosing-regimen, delivery-device, and process claims, which is what a dense peptide thicket demands. The ontology clusters the thicket by concept and normalizes assignees, so a team sees the structure of protection around a molecule rather than a flat list. Cypris Q, the platform's agentic layer, lets teams run and chain FTO analysis conversationally, and Agentic Monitoring tracks a molecule and its surrounding thicket over time, flagging new formulation, device, and combination filings as they publish. Cypris provides enterprise API partnerships with OpenAI, Anthropic, and Google, and is built with enterprise-grade security. Cypris serves hundreds of enterprise customers across pharmaceuticals, chemicals, advanced materials, energy, and other regulated industries.
FAQ
Why is freedom-to-operate hard for GLP-1 and peptide therapeutics?
Freedom-to-operate is hard for GLP-1 and peptide therapeutics because protection is built as a dense, layered thicket extending well beyond the active ingredient. A peer-reviewed analysis found GLP-1 products carry a median of 19.5 listed patents each, most of them on delivery devices. Assessing FTO requires reading composition, formulation, dosing, device, and process claims across many patents and markets.
What claim types create FTO risk for GLP-1 products?
Five claim types create FTO risk for GLP-1 products: composition-of-matter claims on the peptide, formulation claims on stabilized and oral forms, dosing-regimen and method-of-use claims, delivery-device claims, and process or manufacturing claims. Each can independently block a product. Delivery-device claims are a particularly large share of the GLP-1 thicket.
How many patents protect a typical GLP-1 product?
A peer-reviewed analysis of GLP-1 receptor agonists approved between 2005 and 2021 found a median of 19.5 listed patents per product, with 54 percent on delivery devices rather than the active ingredient, and a median of 18.3 years of expected protection after approval. No generic manufacturer had successfully challenged a GLP-1 receptor agonist patent as of that analysis. These figures illustrate the density of the thicket.
What is a pharmaceutical patent thicket?
A pharmaceutical patent thicket is a dense set of overlapping patents around a single product that extends protection beyond the core molecule. Empirical analysis shows patents per active ingredient rose from 1.86 in 2001 to nearly six by 2019, driven substantially by continuation applications. Thickets shape when and where competition can enter.
How does the expiry of GLP-1 patents affect freedom-to-operate?
The expiry of GLP-1 patents affects freedom-to-operate market-by-market, because primary-patent expiry timing varies by jurisdiction and the practical entry date is governed by the surrounding formulation, device, and process claims rather than the molecule alone. FTO must therefore be assessed claim-by-claim and market-by-market. A molecule can be off-patent in one country and still protected in another.
Where is the white space in GLP-1 and peptide development?
Recent filing activity has shifted toward oral delivery, dual and triple receptor agonists, and combination therapies, which is where both new FTO risk and remaining white space now sit. Mapping this frontier requires reading patents and scientific literature together, since peptide innovation appears in research first. White space analysis identifies the areas that are still open.
How does AI-powered FTO help with peptide therapeutics?
AI-powered FTO helps with peptide therapeutics by using semantic search to retrieve relevant claims regardless of terminology, claim-level analysis to focus on the independent claims that define infringement across all claim types, and continuous monitoring to keep a cleared position current. This is what a dense, fast-moving thicket requires. Cypris runs this across more than 500 million patents and scientific papers.
Which teams need GLP-1 and peptide FTO analysis?
GLP-1 and peptide FTO analysis is needed by pharmaceutical and biotech R&D, IP, and business-development teams developing follow-on peptides, biosimilars, differentiated formulations, or combination products. It is also relevant to generics manufacturers assessing entry. Cypris serves hundreds of enterprise customers across pharmaceuticals and other regulated industries.
How current does GLP-1 FTO need to be?
GLP-1 FTO needs to be continuously current, because new formulation, device, dosing, and combination filings publish constantly and can change a cleared position. A one-time assessment reflects only the moment it was run. Cypris uses Agentic Monitoring to track a molecule and its surrounding thicket over time and flag new filings as they publish.
Endnotes
- Tu, S. S., Feldman, W. B., Alhiary, R., Gabriele, S., Kesselheim, A. S. & Beall, R. F. (2023). Patents and Regulatory Exclusivities on GLP-1 Receptor Agonists. JAMA. https://doi.org/10.1001/jama.2023.13872
- Industry analyst estimates (e.g., Research and Markets; BCC Research). GLP-1 market forecasts vary widely by scope and are presented here as order-of-magnitude estimates, not authoritative figures.
- Han, J., Zhou, Z., Jiang, N. & Lu, W. (2023). An updated patent review of GLP-1 receptor agonists (2020–present). Expert Opinion on Therapeutic Patents. https://doi.org/10.1080/13543776.2023.2274905
- Tu, S. S., Feldman, W. B. et al. (2024). Delivery Device Patents on GLP-1 Receptor Agonists. JAMA. https://doi.org/10.1001/jama.2024.0919
- Fasi, M. A. (2026). Patent landscape and therapeutic evolution of mazdutide. Expert Opinion on Therapeutic Patents. https://doi.org/10.1080/13543776.2026.2645812
- Tu, S. S. (2024). The Long CON: An Empirical Analysis of Pharmaceutical Patent Thickets. University of Pittsburgh Law Review. https://doi.org/10.5195/lawreview.2024.1049
- Hemphill, C. S. & Sampat, B. N. (2012). Evergreening, patent challenges, and effective market life in pharmaceuticals. Journal of Health Economics. https://doi.org/10.1016/j.jhealeco.2012.01.004
- Tu, S. S. & Carrier, M. A. (2023). Why Pharmaceutical Patent Thickets Are Unique. SSRN. https://doi.org/10.2139/ssrn.4571486
- Ramesh, S., Cross, S., Levi, J., Hill, A. & Venter, F. (2026). How Low Could Semaglutide Prices Fall? Implications for Global Access Ahead of Patent Expiry. Obesity. https://doi.org/10.1002/oby.70241
Reports
Webinars

Many enterprises have adopted horizontal, foundation-model AI platforms. But access to the same underlying models does not, by itself, create differentiated intelligence. For highly technical and mission-critical research, general-purpose models may produce broad but weakly grounded answers when they lack access to authoritative technical data, specialized context, and verifiable sources.
The next competitive advantage will come from the intelligence layer surrounding the foundation model: the domain-specific data, ontologies, retrieval capabilities, agent workflows, and source grounding that together form an AI harness. These verticalized systems can transform general-purpose AI into a more specialized capability for research, innovation, and technical decision-making.
Join Steve Hafif, Co-Founder and CEO of Cypris.ai, and Marlene Valderrama, Principal IP Manager and Senior Technology Scout at Halliburton, for a conversation on the state of enterprise AI and how organizations can enhance horizontal AI platforms with verticalized intelligence designed for R&D and innovation.
.png)

Most IP organizations are making high-stakes capital allocation decisions with incomplete visibility – relying primarily on patent data as a proxy for innovation. That approach is not optimal. Patents alone cannot reveal technology trajectories, capital flows, or commercial viability.
A more effective model requires integrating patents with scientific literature, grant funding, market activity, and competitive intelligence. This means that for a complete picture, IP and R&D teams need infrastructure that connects fragmented data into a unified, decision-ready intelligence layer.
AI is accelerating that shift. The value is no longer simply in retrieving documents faster; it’s in extracting signal from noise. Modern AI systems can contextualize disparate datasets, identify patterns, and generate strategic narratives – transforming raw information into actionable insight.
Join us on Thursday, April 23, at 12 PM ET for a discussion on how unified AI platforms are redefining decision-making across IP and R&D teams. Moderated by Gene Quinn, panelists Marlene Valderrama and Amir Achourie will examine how integrating technical, scientific, and market data collapses traditional silos – enabling more aligned strategy, sharper investment decisions, and measurable business impact.
Register here: https://ipwatchdog.com/cypris-april-23-2026/
.png)
In this session, we break down how AI is reshaping the R&D lifecycle, from faster discovery to more informed decision-making. See how an intelligence layer approach enables teams to move beyond fragmented tools toward a unified, scalable system for innovation.
.avif)

%20-%20Competitive%20Benchmarking%20for%20Industrial%20Robotics%20OEMs.png)
%20-%20Competitive%20Benchmarking%20of%20EV%20Battery%20Material%20%26%20Cell%20Manufacturers.png)
%20-%20Competitive%20Benchmarking%20for%20Wearable%20%26%20Biosensor%20Device%20Manufacturers.png)