
Insights on Innovation, R&D, and IP
Perspectives on patents, scientific research, emerging technologies, and the strategies shaping modern R&D

Executive Summary
In 2024, US patent infringement jury verdicts totaled $4.19 billion across 72 cases. Twelve individual verdicts exceeded $100million. The largest single award—$857 million in General Access Solutions v.Cellco Partnership (Verizon)—exceeded the annual R&D budget of many mid-market technology companies. In the first half of 2025 alone, total damages reached an additional $1.91 billion.
The consequences of incomplete patent intelligence are not abstract. In what has become one of the most instructive IP disputes in recent history, Masimo’s pulse oximetry patents triggered a US import ban on certain Apple Watch models, forcing Apple to disable its blood oxygen feature across an entire product line, halt domestic sales of affected models, invest in a hardware redesign, and ultimately face a $634 million jury verdict in November 2025. Apple—a company with one of the most sophisticated intellectual property organizations on earth—spent years in litigation over technology it might have designed around during development.
For organizations with fewer resources than Apple, the risk calculus is starker. A mid-size materials company, a university spinout, or a defense contractor developing next-generation battery technology cannot absorb a nine-figure verdict or a multi-year injunction. For these organizations, the patent landscape analysis conducted during the development phase is the primary risk mitigation mechanism. The quality of that analysis is not a matter of convenience. It is a matter of survival.
And yet, a growing number of R&D and IP teams are conducting that analysis using general-purpose AI tools—ChatGPT, Claude, Microsoft Co-Pilot—that were never designed for patent intelligence and are structurally incapable of delivering it.
This report presents the findings of a controlled comparison study in which identical patent landscape queries were submitted to four AI-powered tools: Cypris (a purpose-built R&D intelligence platform),ChatGPT (OpenAI), Claude (Anthropic), and Microsoft Co-Pilot. Two technology domains were tested: solid-state lithium-sulfur battery electrolytes using garnet-type LLZO ceramic materials (freedom-to-operate analysis), and bio-based polyamide synthesis from castor oil derivatives (competitive intelligence).
The results reveal a significant and structurally persistent gap. In Test 1, Cypris identified over 40 active US patents and published applications with granular FTO risk assessments. Claude identified 12. ChatGPT identified 7, several with fabricated attribution. Co-Pilot identified 4. Among the patents surfaced exclusively by Cypris were filings rated as “Very High” FTO risk that directly claim the technology architecture described in the query. In Test 2, Cypris cited over 100 individual patent filings with full attribution to substantiate its competitive landscape rankings. No general-purpose model cited a single patent number.
The most active sectors for patent enforcement—semiconductors, AI, biopharma, and advanced materials—are the same sectors where R&D teams are most likely to adopt AI tools for intelligence workflows. The findings of this report have direct implications for any organization using general-purpose AI to inform patent strategy, competitive intelligence, or R&D investment decisions.

1. Methodology
A controlled comparative evaluation was conducted on March 27, 2026. An identical patent landscape query was submitted verbatim to each platform under standardized testing conditions. No follow-up prompts, clarifications, or iterative refinements were permitted, ensuring that each platform was evaluated based solely on its initial response.
The outputs were preserved in their original form and evaluated against predefined criteria using publicly verifiable patent records.
1.1 Query
Identify all active US patents and published applications filed in the last 5 years related to solid-state lithium-sulfur battery electrolytes using garnet-type ceramic materials. For each, provide the assignee, filing date, key claims, and current legal status. Highlight any patents that could pose freedom-to-operate risks for a company developing a Li₇La₃Zr₂O₁₂(LLZO)-based composite electrolyte with a polymer interlayer.
1.2 Tools Evaluated

1.3 Evaluation Criteria
Each response was evaluated using a consistent six-part scoring framework: patent coverage, assignee accuracy, filing metadata completeness, depth of claim analysis, quality of FTO risk stratification, and the presence of actionable strategic guidance.
Patent numbers, assignees, filing information, and legal status were independently checked against publicly available USPTO and WIPO records. The evaluation focused on the completeness, accuracy, and practical utility of each platform’s output rather than writing quality or presentation.
2. Findings
2.1 Coverage Gap
The most significant finding is the scale of the coverage differential. Cypris identified over 40 active US patents and published applications spanning LLZO-polymer composite electrolytes, garnet interface modification, polymer interlayer architectures, lithium-sulfur specific filings, and adjacent ceramic composite patents. The results were organized by technology category with per-patent FTO risk ratings.
Claude identified 12 patents organized in a four-tier risk framework. Its analysis was structurally sound and correctly flagged the two highest-risk filings (Solid Energies US 11,967,678 and the LLZO nanofiber multilayer US 11,923,501). It also identified the University ofMaryland/ Wachsman portfolio as a concentration risk and noted the NASA SABERS portfolio as a licensing opportunity. However, it missed the majority of the landscape, including the entire Corning portfolio, GM's interlayer patents, theKorea Institute of Energy Research three-layer architecture, and the HonHai/SolidEdge lithium-sulfur specific filing.
ChatGPT identified 7 patents, but the quality of attribution was inconsistent. It listed assignees as "Likely DOE /national lab ecosystem" and "Likely startup / defense contractor cluster" for two filings—language that indicates the model was inferring rather than retrieving assignee data. In a freedom-to-operate context, an unverified assignee attribution is functionally equivalent to no attribution, as it cannot support a licensing inquiry or risk assessment.
Co-Pilot identified 4 US patents. Its output was the most limited in scope, missing the Solid Energies portfolio entirely, theUMD/ Wachsman portfolio, Gelion/ Johnson Matthey, NASA SABERS, and all Li-S specific LLZO filings.
2.2 Critical Patents Missed by Public Models
The following table presents patents identified exclusively by Cypris that were rated as High or Very High FTO risk for the proposed technology architecture. None were surfaced by any general-purpose model.

2.3 Patent Fencing: The Solid Energies Portfolio
Cypris identified a coordinated patent fencing strategy by Solid Energies, Inc. that no general-purpose model detected at scale. Solid Energies holds at least four granted US patents and one published application covering LLZO-polymer composite electrolytes across compositions(US-12463245-B2), gradient architectures (US-12283655-B2), electrode integration (US-12463249-B2), and manufacturing processes (US-20230035720-A1). Claude identified one Solid Energies patent (US 11,967,678) and correctly rated it as the highest-priority FTO concern but did not surface the broader portfolio. ChatGPT and Co-Pilot identified zero Solid Energies filings.
The practical significance is that a company relying on any individual patent hit would underestimate the scope of Solid Energies' IP position. The fencing strategy—covering the composition, the architecture, the electrode integration, and the manufacturing method—means that identifying a single design-around for one patent does not resolve the FTO exposure from the portfolio as a whole. This is the kind of strategic insight that requires seeing the full picture, which no general-purpose model delivered
2.4 Assignee Attribution Quality
ChatGPT's response included at least two instances of fabricated or unverifiable assignee attributions. For US 11,367,895 B1, the listed assignee was "Likely startup / defense contractor cluster." For US 2021/0202983 A1, the assignee was described as "Likely DOE / national lab ecosystem." In both cases, the model appears to have inferred the assignee from contextual patterns in its training data rather than retrieving the information from patent records.
In any operational IP workflow, assignee identity is foundational. It determines licensing strategy, litigation risk, and competitive positioning. A fabricated assignee is more dangerous than a missing one because it creates an illusion of completeness that discourages further investigation. An R&D team receiving this output might reasonably conclude that the landscape analysis is finished when it is not.
3. Structural Limitations of General-Purpose Models for Patent Intelligence
3.1 Training Data Is Not Patent Data
Large language models are trained on web-scraped text. Their knowledge of the patent record is derived from whatever fragments appeared in their training corpus: blog posts mentioning filings, news articles about litigation, snippets of Google Patents pages that were crawlable at the time of data collection. They do not have systematic, structured access to the USPTO database. They cannot query patent classification codes, parse claim language against a specific technology architecture, or verify whether a patent has been assigned, abandoned, or subjected to terminal disclaimer since their training data was collected.
This is not a limitation that improves with scale. A larger training corpus does not produce systematic patent coverage; it produces a larger but still arbitrary sampling of the patent record. The result is that general-purpose models will consistently surface well-known patents from heavily discussed assignees (QuantumScape, for example, appeared in most responses) while missing commercially significant filings from less publicly visible entities (Solid Energies, Korea Institute of EnergyResearch, Shenzhen Solid Advanced Materials).
3.2 The Web Is Closing to Model Scrapers
The data access problem is structural and worsening. As of mid-2025, Cloudflare reported that among the top 10,000 web domains, the majority now fully disallow AI crawlers such as GPTBot andClaudeBot via robots.txt. The trend has accelerated from partial restrictions to outright blocks, and the crawl-to-referral ratios reveal the underlying tension: OpenAI's crawlers access approximately1,700 pages for every referral they return to publishers; Anthropic's ratio exceeds 73,000 to 1.
Patent databases, scientific publishers, and IP analytics platforms are among the most restrictive content categories. A Duke University study in 2025 found that several categories of AI-related crawlers never request robots.txt files at all. The practical consequence is that the knowledge gap between what a general-purpose model "knows" about the patent landscape and what actually exists in the patent record is widening with each training cycle. A landscape query that a general-purpose model partially answered in 2023 may return less useful information in 2026.
3.3 General-Purpose Models Lack Ontological Frameworks for Patent Analysis
A freedom-to-operate analysis is not a summarization task. It requires understanding claim scope, prosecution history, continuation and divisional chains, assignee normalization (a single company may appear under multiple entity names across patent records), priority dates versus filing dates versus publication dates, and the relationship between dependent and independent claims. It requires mapping the specific technical features of a proposed product against independent claim language—not keyword matching.
General-purpose models do not have these frameworks. They pattern-match against training data and produce outputs that adopt the format and tone of patent analysis without the underlying data infrastructure. The format is correct. The confidence is high. The coverage is incomplete in ways that are not visible to the user.
4. Comparative Output Quality
The following table summarizes the qualitative characteristics of each tool's response across the dimensions most relevant to an operational IP workflow.

5. Implications for R&D and IP Organizations
5.1 The Confidence Problem
The central risk identified by this study is not that general-purpose models produce bad outputs—it is that they produce incomplete outputs with high confidence. Each model delivered its results in a professional format with structured analysis, risk ratings, and strategic recommendations. At no point did any model indicate the boundaries of its knowledge or flag that its results represented a fraction of the available patent record. A practitioner receiving one of these outputs would have no signal that the analysis was incomplete unless they independently validated it against a comprehensive datasource.
This creates an asymmetric risk profile: the better the format and tone of the output, the less likely the user is to question its completeness. In a corporate environment where AI outputs are increasingly treated as first-pass analysis, this dynamic incentivizes under-investigation at precisely the moment when thoroughness is most critical.
5.2 The Diversification Illusion
It might be assumed that running the same query through multiple general-purpose models provides validation through diversity of sources. This study suggests otherwise. While the four tools returned different subsets of patents, all operated under the same structural constraints: training data rather than live patent databases, web-scraped content rather than structured IP records, and general-purpose reasoning rather than patent-specific ontological frameworks. Running the same query through three constrained tools does not produce triangulation; it produces three partial views of the same incomplete picture.
5.3 The Appropriate Use Boundary
General-purpose language models are effective tools for a wide range of tasks: drafting communications, summarizing documents, generating code, and exploratory research. The finding of this study is not that these tools lack value but that their value boundary does not extend to decisions that carry existential commercial risk.
Patent landscape analysis, freedom-to-operate assessment, and competitive intelligence that informs R&D investment decisions fall outside that boundary. These are workflows where the completeness and verifiability of the underlying data are not merely desirable but are the primary determinant of whether the analysis has value. A patent landscape that captures 10% of the relevant filings, regardless of how well-formatted or confidently presented, is a liability rather than an asset.
6. Test 2: Competitive Intelligence — Bio-Based Polyamide Patent Landscape
To assess whether the findings from Test 1 were specific to a single technology domain or reflected a broader structural pattern, a second query was submitted to all four tools. This query shifted from freedom-to-operate analysis to competitive intelligence, asking each tool to identify the top 10organizations by patent filing volume in bio-based polyamide synthesis from castor oil derivatives over the past three years, with summaries of technical approach, co-assignee relationships, and portfolio trajectory.
6.1 Query

6.2 Summary of Results

6.3 Key Differentiators
Verifiability
The most consequential difference in Test 2 was the presence or absence of verifiable evidence. Cypris cited over 100 individual patent filings with full patent numbers, assignee names, and publication dates. Every claim about an organization’s technical focus, co-assignee relationships, and filing trajectory was anchored to specific documents that a practitioner could independently verify in USPTO, Espacenet, or WIPO PATENT SCOPE. No general-purpose model cited a single patent number. Claude produced the most structured and analytically useful output among the public models, with estimated filing ranges, product names, and strategic observations that were directionally plausible. However, without underlying patent citations, every claim in the response requires independent verification before it can inform a business decision. ChatGPT and Co-Pilot offered thinner profiles with no filing counts and no patent-level specificity.
Data Integrity
ChatGPT’s response contained a structural error that would mislead a practitioner: it listed CathayBiotech as organization #5 and then listed “Cathay Affiliate Cluster” as a separate organization at #9, effectively double-counting a single entity. It repeated this pattern with Toray at #4 and “Toray(Additional Programs)” at #10. In a competitive intelligence context where the ranking itself is the deliverable, this kind of error distorts the landscape and could lead to misallocation of competitive monitoring resources.
Organizations Missed
Cypris identified Kingfa Sci. & Tech. (8–10 filings with a differentiated furan diacid-based polyamide platform) and Zhejiang NHU (4–6 filings focused on continuous polymerization process technology)as emerging players that no general-purpose model surfaced. Both represent potential competitive threats or partnership opportunities that would be invisible to a team relying on public AI tools.Conversely, ChatGPT included organizations such as ANTA and Jiangsu Taiji that appear to be downstream users rather than significant patent filers in synthesis, suggesting the model was conflating commercial activity with IP activity.
Strategic Depth
Cypris’s cross-cutting observations identified a fundamental chemistry divergence in the landscape:European incumbents (Arkema, Evonik, EMS) rely on traditional castor oil pyrolysis to 11-aminoundecanoic acid or sebacic acid, while Chinese entrants (Cathay Biotech, Kingfa) are developing alternative bio-based routes through fermentation and furandicarboxylic acid chemistry.This represents a potential long-term disruption to the castor oil supply chain dependency thatWestern players have built their IP strategies around. Claude identified a similar theme at a higher level of abstraction. Neither ChatGPT nor Co-Pilot noted the divergence.
6.4 Test 2 Conclusion
Test 2 confirms that the coverage and verifiability gaps observed in Test 1 are not domain-specific.In a competitive intelligence context—where the deliverable is a ranked landscape of organizationalIP activity—the same structural limitations apply. General-purpose models can produce plausible-looking top-10 lists with reasonable organizational names, but they cannot anchor those lists to verifiable patent data, they cannot provide precise filing volumes, and they cannot identify emerging players whose patent activity is visible in structured databases but absent from the web-scraped content that general-purpose models rely on.
7. Conclusion
This comparative analysis, spanning two distinct technology domains and two distinct analytical workflows—freedom-to-operate assessment and competitive intelligence—demonstrates that the gap between purpose-built R&D intelligence platforms and general-purpose language models is not marginal, not domain-specific, and not transient. It is structural and consequential.
In Test 1 (LLZO garnet electrolytes for Li-S batteries), the purpose-built platform identified more than three times as many patents as the best-performing general-purpose model and ten times as many as the lowest-performing one. Among the patents identified exclusively by the purpose-built platform were filings rated as Very High FTO risk that directly claim the proposed technology architecture. InTest 2 (bio-based polyamide competitive landscape), the purpose-built platform cited over 100individual patent filings to substantiate its organizational rankings; no general-purpose model cited as ingle patent number.
The structural drivers of this gap—reliance on training data rather than live patent feeds, the accelerating closure of web content to AI scrapers, and the absence of patent-specific analytical frameworks—are not transient. They are inherent to the architecture of general-purpose models and will persist regardless of increases in model capability or training data volume.
For R&D and IP leaders, the practical implication is clear: general-purpose AI tools should be used for general-purpose tasks. Patent intelligence, competitive landscaping, and freedom-to-operate analysis require purpose-built systems with direct access to structured patent data, domain-specific analytical frameworks, and the ability to surface what a general-purpose model cannot—not because it chooses not to, but because it structurally cannot access the data.
The question for every organization making R&D investment decisions today is whether the tools informing those decisions have access to the evidence base those decisions require. This study suggests that for the majority of general-purpose AI tools currently in use, the answer is no.
Study Disclosure
This comparative evaluation was commissioned and published by Cypris. The testing methodology, prompts, evaluation criteria, and underlying outputs have been documented to support independent review and replication.
All platform outputs were preserved in their original form. Patent data and material factual claims were cross-checked against USPTO Patent Center and WIPO PATENTSCOPE records as of March 27, 2026. Cypris was one of the platforms evaluated and therefore has a commercial interest in the findings.
The Patent Intelligence Gap - A Comparative Analysis of Verticalized AI-Patent Tools vs. General-Purpose Language Models for R&D Decision-Making
All Blogs

Chemical and enzymatic recycling has become the frontier of the circular economy for plastics, and its patent landscape is distinctive because molecular recycling is not one technology but a set of competing routes, each with its own chemistry, feedstocks, and process engineering. Where mechanical recycling melts and reforms plastic, losing quality with each cycle and struggling with colored, mixed, or contaminated waste, molecular recycling breaks polymers back down to their building blocks, monomers or feedstock chemicals, that can be repurified and repolymerized to a quality equivalent to virgin material. The routes divide by polymer and mechanism: engineered enzymes depolymerize polyester and PET even when colored or contaminated, with machine-learning-aided enzyme engineering producing fast, robust variants;¹ chemical routes such as methanolysis, glycolysis, and hydrolysis cleave PET into its monomers; and pyrolysis and gasification convert polyolefins such as polyethylene and polypropylene into oils and feedstocks. A persistent constraint on the enzymatic route is that highly crystalline PET resists depolymerization, so pretreatment and reaction-medium engineering matter as much as the enzyme itself,³ and thermostable, efficient enzymes remain difficult to engineer.⁴,⁵ Because each route is a distinct region of patenting, freedom-to-operate and white space analysis must treat plastic recycling as several landscapes at once, spanning enzymes and catalysts, reactor and process design, feedstock pretreatment, and monomer purification.
The field is being pulled forward by regulation and by the arrival of commercial-scale plants. In the European Union, the Packaging and Packaging Waste Regulation sets binding minimum-recycled-content targets for plastic packaging for 2030 and 2040 and requires packaging to be recyclable by 2030, creating durable demand for high-quality recycled material that mechanical recycling cannot fully supply.⁶ At the same time, enzyme discovery is industrializing: a systematic profiling of PET-depolymerizing enzymes mapped roughly 1,894 candidate hydrolases across about 170 natural sequence lineages, vastly expanding the toolbox beyond the handful of enzymes studied a few years ago.² Commercial facilities are moving from demonstration to build-out, with a first industrial enzymatic PET plant designed for about 50,000 tonnes per year of prepared post-consumer waste, on a revised commissioning schedule.⁷ The patent record reflects both the biology and the chemistry: across the Cypris corpus of more than 500 million patents and scientific papers, the PET-depolymerization and molecular-recycling set holds on the order of 1,266 families and grew from about 26 in 2020 to roughly 171 in 2024, with the most active assignees mixing chemical-route incumbents such as IFP Energies Nouvelles and Eastman with enzymatic and brand-side players such as Jeplan and Coca-Cola, and the United States, France, and China leading on geography; 2025 and 2026 counts are partial because of the publication lag.
The strategic question is which route and polymer to back, and the white space sits where the chemistry is hardest. Engineered enzymes have advanced furthest for PET and polyester, so the open, high-value ground is increasingly in enzymes and processes for other polymers, in the machine-learning-guided discovery and engineering of new depolymerizing enzymes,²,⁵ and in handling mixed and contaminated feedstocks. Chemical routes for PET are maturing, while the chemical recycling of polyolefins, the largest share of plastic waste, remains harder and less crowded, particularly in upgrading pyrolysis oils to usable feedstocks. Textile-to-textile recycling of polyester is a further emerging layer. Reading the landscape by route, polymer, and process, and tracking both the patents and the underlying enzyme and catalysis research, is what separates a crowded region from an open one.
Where the plastic-recycling white space is
Enzymes for non-PET polymers. Engineered enzymes that depolymerize polyolefins, nylons, and polyurethanes, rather than only polyester, are an early, high-value target as enzymatic PET matures.
Machine-learning enzyme discovery. Computational discovery and engineering of new, more thermostable and efficient depolymerases is a fast-moving layer that compresses development time.²,⁵
Polyolefin chemical recycling. Pyrolysis and its upgrading to usable feedstocks for the largest category of plastic waste remain harder and less crowded than PET routes.
Mixed and contaminated feedstocks. Processes that handle colored, multilayer, and mixed waste, which mechanical recycling cannot, are a differentiating capability.
Monomer purification and textile recycling. High-purity monomer recovery and textile-to-textile polyester recycling are distinct, emerging layers where quality and economics are decided.
How AI-powered landscape and white space analysis helps
Resolving a landscape that spans enzymatic and chemical routes across several polymers, under a regulatory recycled-content pull, requires more than keyword search. AI-powered analysis addresses this with semantic search that clusters activity by route, polymer, and process across varied terminology, attribution that normalizes filers to canonical entities and tracks new entrants, and continuous monitoring that keeps pace with a fast-commercializing field. Because recycling advances appear in scientific and enzyme-engineering literature before they are patented, reading both patents and literature gives the earliest signal of where scalable routes are emerging.
Where Cypris fits
Cypris runs patent landscape and white space analysis for multi-route fields such as chemical and enzymatic plastic recycling across a corpus of more than 500 million patents and scientific papers, organized through a proprietary R&D ontology. The ontology clusters activity by route, enzymatic, methanolysis, glycolysis, hydrolysis, and pyrolysis, and by polymer and process layer, and normalizes filers to canonical entities, so a team can resolve which routes and polymers are crowded and which remain open as white space, and can track new entrants as the field scales. Semantic search across patents and scientific literature connects filings to the underlying enzyme-engineering and catalysis research, which is where recycling advances appear first. Cypris Q, the platform's agentic layer, lets teams run landscape and white space analysis conversationally and chain the clustering, attribution, and gap analysis, and Agentic Monitoring tracks a defined route over time and flags new patents and papers as they publish. Cypris provides enterprise API partnerships with OpenAI, Anthropic, and Google, and is built with enterprise-grade security. Cypris serves hundreds of enterprise customers across pharmaceuticals, chemicals, advanced materials, energy, and other regulated industries.
FAQ
What is chemical and enzymatic plastic recycling? Chemical and enzymatic plastic recycling, also called molecular recycling, breaks plastics back down to monomers or feedstock chemicals that can be repolymerized to virgin quality, unlike mechanical recycling, which degrades quality. Routes include engineered-enzyme depolymerization, methanolysis, glycolysis, and hydrolysis for PET, and pyrolysis and gasification for polyolefins. Each is a distinct region of patenting.
Why is molecular recycling a patenting hotspot? Molecular recycling is a patenting hotspot because recycled-content rules and packaging regulations are creating durable demand for high-quality recycled material, and the first commercial plants are coming online. It can process colored, mixed, and contaminated waste that mechanical recycling cannot. That combination is driving filings across enzymes, catalysts, and processes.
What routes does the plastic-recycling landscape cover? The landscape covers engineered-enzyme depolymerization of polyester and PET, chemical depolymerization of PET by methanolysis, glycolysis, and hydrolysis, and pyrolysis and gasification of polyolefins. Each route has distinct enzymes or catalysts, reactors, and purification steps. Freedom-to-operate and white space analysis must treat them separately.
Where is the white space in plastic recycling? The white space includes engineered enzymes for non-PET polymers, machine-learning-guided enzyme discovery, polyolefin chemical recycling and pyrolysis-oil upgrading, processes for mixed and contaminated feedstocks, and monomer purification and textile-to-textile recycling. Enzymatic PET is comparatively advanced. The most open, high-value opportunities are in other polymers and in polyolefin routes.
Why is polyolefin recycling harder than PET recycling? Polyolefin recycling is harder because polyethylene and polypropylene lack the cleavable bonds that make PET amenable to enzymatic and chemical depolymerization, so they are typically broken down by pyrolysis into mixed oils that must then be upgraded. Polyolefins are also the largest share of plastic waste. That difficulty leaves the layer less crowded and high in value.
Why does plastic-recycling analysis need scientific literature? Plastic-recycling analysis needs scientific literature because enzyme-engineering and catalysis advances appear in research before they are patented, so the literature gives the earliest signal. Analyzing patents alone gives a lagging view. Cypris analyzes both across more than 500 million patents and scientific papers.
What software helps analyze the plastic-recycling patent landscape? Software for the plastic-recycling landscape should cluster activity by route, polymer, and process, resolve filers to canonical owners, search patents and scientific literature semantically, and monitor a fast-commercializing field continuously. Cypris does this across more than 500 million patents and scientific papers using a proprietary R&D ontology, semantic search, Cypris Q, and Agentic Monitoring.
Which teams use plastic-recycling patent landscape analysis? Plastic-recycling patent landscape analysis is used by R&D, innovation, IP, and strategy teams at chemicals, materials, packaging, and waste-management companies, enzyme and catalysis developers, and their partners, as well as investors and policymakers. It informs which route to back, where to file, and where competitors are concentrated. Cypris serves hundreds of enterprise customers across chemicals, advanced materials, energy, and other regulated industries.
Endnotes
- Lu, H., Diaz, D. J., Shroff, P., Alper, H. S., et al. (2022). Machine learning-aided engineering of hydrolases for PET depolymerization. Nature, 604. https://doi.org/10.1038/s41586-022-04599-z
- Park, S. Y., Ki, D., Sagong, H.-Y., Seo, H., et al. (2025). Landscape profiling of PET depolymerases using a natural sequence cluster framework. Science, 387(6729). https://doi.org/10.1126/science.adp5637
- Kaabel, S., Therien, J. P. D., Auclair, K., et al. (2021). Enzymatic depolymerization of highly crystalline PET enabled in moist-solid reaction mixtures. Proceedings of the National Academy of Sciences, 118(29). https://doi.org/10.1073/pnas.2026452118
- Kawabata, T., Iizuka, R., & Kawai, F. (2024). Engineered polyethylene terephthalate hydrolases: perspectives and limits. Applied Microbiology and Biotechnology, 108. https://doi.org/10.1007/s00253-024-13222-2
- Li, Q., Su, L., Dian, L., et al. (2024). Dynamic docking-assisted engineering of hydrolases for efficient PET depolymerization. ACS Catalysis, 14(9). https://doi.org/10.1021/acscatal.4c00400
- European Commission, Directorate-General for Environment. Packaging and packaging waste. https://environment.ec.europa.eu/topics/waste-and-recycling/packaging-waste_en
- Carbios. PET biorecycling technology. https://www.carbios.com/en/pet-biorecycling-technology/

The fastest way to turn a commodity AI assistant into a reliable R&D and IP research tool is to connect it to a domain-oriented intelligence layer through the Model Context Protocol, because the general-purpose model supplies the reasoning while the verticalized agent supplies the grounded, high-signal data the model cannot hold on its own. This is the single architectural decision that separates an AI that drafts plausible-sounding patent summaries from one an innovation team can actually act on. The model you start with is a commodity. The vertical integration you attach to it is the differentiator.
This guide explains what commodity AI gets wrong in R&D and IP work, why the gap is structural rather than a matter of prompting, and how a domain MCP integration closes it. It is written for R&D directors, IP managers, and innovation strategists who already have access to capable general models and want to understand what it takes to make them trustworthy for stage-gate decisions.
What Commodity AI Means in an R&D Context
A commodity AI is a general-purpose large language model accessed through a chat interface or an enterprise assistant, the same model available to every competitor in your market. These horizontal systems are built on broad pre-training across diverse public data and are designed to handle a wide range of tasks without deep subject knowledge [1]. They are genuinely useful for summarizing a document you paste in, drafting an email, or explaining a concept. The strength of the horizontal model is breadth and speed of deployment.
The weakness is that breadth is the wrong shape for R&D and IP intelligence. A prior art search, a freedom-to-operate question, or a white space analysis does not reward general fluency. It rewards completeness, recency, and precision against a defined corpus of patents and scientific literature. A commodity model has no live connection to that corpus. It answers from a frozen snapshot of training data and from whatever you happened to paste into the prompt, which means the most consequential R&D questions are exactly the ones it is least equipped to answer.
Why the Gap Is Structural, Not a Prompting Problem
The instinct when a general model gives a weak patent answer is to write a better prompt. This helps at the margin, but it cannot solve the core problem, because the failure is rooted in two structural limits that prompting does not touch.
The first limit is hallucination. Generating plausible but ungrounded output remains the single biggest barrier to deploying language models in production as of 2026, and complete elimination is not possible because the tendency is tied to the model's generative capability itself [2]. In an IP context this is not a cosmetic flaw. A model conducting an ungrounded prior art search can surface references that do not exist, misattribute a claim, or describe a system that is physically impossible, and it delivers all of it in the same confident register as a correct answer [3]. A 2026 study evaluating five popular public models on preliminary prior art searches found that accuracy, consistency, and the ability to surface conceptually relevant art from adjacent fields varied widely and required careful human verification [4]. The authority of the output is not evidence of its reliability.
The second limit is that flooding a general model with more data does not fix the first problem and often makes it worse. There is a temptation to solve grounding by dumping an entire patent dataset into the model's context window. Research on context engineering shows this backfires. As a broad, undifferentiated corpus fills the context window, the model's ability to reason over it degrades, an effect documented across multiple studies of how models use long contexts [5][6]. The model does not get smarter as you add data. Past a point, it gets less accurate. This is why raw access to a large dataset is not the same as intelligence over it, and why the path to reliability runs through retrieving the right small set of high-signal documents rather than the largest possible set.
Together these two limits define the gap. The commodity model is fluent but ungrounded, and you cannot ground it simply by giving it everything. You ground it by connecting it to a system that already knows which fraction of the corpus matters for the question being asked.
What a Verticalized Agent Adds
A vertical AI agent is purpose-built for a specific domain, pre-loaded with domain knowledge, proprietary data models, and deep integrations into the systems where that domain's data lives [7]. Where a horizontal agent relies on broad pre-training, a vertical agent demands domain adaptation and plugs into domain-specific data pipelines, and it is this depth that produces superior accuracy, compliance, and reliability within its field [1]. The market has moved decisively in this direction. Industry analysts forecast that vertical-first deployments will account for a large and growing share of enterprise AI in 2026, with industry-specific AI solutions growing far faster than general-purpose tools, because the highest-return deployments come from embedding agents into existing domain workflows rather than buying a generic assistant [8].
In R&D and IP, the domain adaptation that matters is an ontology. A proprietary R&D ontology lets a vertical agent understand that a query about a polymer coating, a thermal barrier, and a specific chemical family are related concepts in a way a keyword search never will, and it lets the agent retrieve the conceptually relevant subset of patents and papers rather than a lexical match. That is the precise capability the commodity model lacks and the precise reason it cannot be prompted into existence. The ontology is the difference between access to 500 million patents and scientific papers and intelligence over them.
Where MCP Fits
The Model Context Protocol is the open standard that lets a general model call an external system as a tool during a conversation, which is what makes the upgrade from commodity AI to verticalized agent a connection rather than a rebuild [9]. You do not have to abandon the general model your team already uses. MCP is the mechanism by which that model reaches out, mid-reasoning, to a domain-oriented layer, asks it a scoped question, and receives back a reasoned, grounded answer rather than a raw dump of records.
This is the architectural pattern that resolves the structural gap. The general model continues to do what it is good at, which is language, synthesis, and conversation. The vertical agent does what it is good at, which is retrieving the high-signal subset from a defined corpus and reasoning within the domain. The protocol connects them. Crucially, because the vertical layer returns a scoped and reasoned result rather than the entire dataset, it sidesteps the context degradation problem entirely. The model never has to hold the full corpus in its context window, so its reasoning stays sharp.
How the Upgrade Works in Practice
The practical sequence is straightforward to describe even though the engineering behind the vertical layer is substantial. A researcher asks a question in the AI interface they already use. The general model recognizes that the question requires domain intelligence and, through MCP, routes a scoped query to the domain-oriented R&D layer. That layer uses its ontology to retrieve the relevant patents and scientific papers, reasons over them within the domain, and returns a grounded finding. The general model then composes that finding into a clear answer for the researcher. The researcher experiences one fluid conversation. Underneath it, the work has been divided between the part of the system built for language and the part built for the domain.
This division maps directly onto the R&D and IP stage-gate process. A prior art agent built this way returns grounded references rather than invented ones. A white space analysis returns a defensible read of where the unclaimed territory sits. A freedom-to-operate question is answered against live patent data rather than a stale training snapshot. Regulatory tracking stays current because the vertical layer, not the frozen model, is the source of truth. In each case the commodity model is the interface and the verticalized agent is the engine.
What This Means for Buyers
The strategic takeaway is that the model is no longer where the advantage lives. Every competitor in your market can access the same capable general models, which is precisely what makes them a commodity. The durable advantage comes from what you connect those models to. An organization that wires its general AI to a domain-oriented R&D intelligence layer through MCP gets grounded, current, defensible answers to its most important innovation questions. An organization that relies on the commodity model alone gets fluent guesses. The gap between those two outcomes is not the model. It is the vertical integration.
Cypris is built to be that vertical layer. As an enterprise R&D intelligence platform spanning more than 500 million patents and scientific papers, organized by a proprietary R&D ontology and powered by Cypris Q agentic workflows, it is designed to deliver domain-oriented intelligence to the AI systems R&D and innovation teams already use, through enterprise API partnerships with OpenAI, Anthropic, and Google [10]. Rather than asking a general model to be an IP expert it cannot be, Cypris supplies the grounded domain reasoning the model needs, across the workflows that matter most: prior art agents, white space analysis, freedom-to-operate, and regulatory tracking. The commodity model handles the conversation. Cypris handles the intelligence.
Frequently Asked Questions
What does it mean to upgrade commodity AI with a vertical agent?
It means connecting a general-purpose AI model to a domain-specific intelligence system so the model can answer specialized questions accurately. The general model provides language and reasoning, while the vertical agent provides grounded, high-signal data from a defined corpus such as patents and scientific papers. The connection is what turns a fluent generalist into a reliable domain tool.
Why can't I just use a better prompt to get good patent answers from a general AI?
Prompting helps at the margin but cannot solve the core problem, because the failure is structural. A general model has no live connection to patent and scientific data and answers from a frozen training snapshot, so it can hallucinate references that do not exist. Better prompts cannot create data access the model fundamentally lacks.
What is the Model Context Protocol and why does it matter here?
The Model Context Protocol, or MCP, is an open standard that lets a general AI model call an external system as a tool during a conversation. It matters because it allows a commodity model to reach a domain-oriented intelligence layer mid-reasoning and receive a grounded answer. MCP is the mechanism that connects a general model to a vertical agent without replacing the model.
Won't connecting my AI to a huge patent database make it smarter?
Not on its own. Research on context engineering shows that flooding a model's context window with a broad, undifferentiated corpus degrades its reasoning rather than improving it. The value comes from a system that retrieves the small, high-signal subset relevant to your question, not from raw access to the largest possible dataset.
What is the difference between a horizontal AI agent and a vertical AI agent?
A horizontal agent is general-purpose and built for breadth across many tasks and departments, with broad pre-training and fast deployment. A vertical agent is purpose-built for a single domain, pre-loaded with domain knowledge and integrated into domain-specific data pipelines. Vertical agents take longer to build but deliver superior accuracy and reliability within their field.
Why is hallucination such a serious problem for R&D and IP work?
Because in prior art and freedom-to-operate work, a confident wrong answer can misdirect a real innovation or legal decision. Hallucination remains the biggest barrier to production deployment of language models in 2026, and a model can surface non-existent references in the same authoritative tone as correct ones. The authority of the output is not evidence of its accuracy.
What role does an ontology play in a vertical R&D agent?
An ontology lets the agent understand conceptual relationships between technologies, materials, and methods rather than relying on keyword matching. This allows it to retrieve patents and papers that are conceptually relevant even when they use different terminology. The ontology is the core capability that makes a vertical agent precise where a general model is not.
Do I have to replace my existing AI tools to do this?
No. The entire point of an MCP-based integration is that you keep the general AI your team already uses and connect it to a vertical intelligence layer. The general model remains the interface, and the domain agent works behind it. The upgrade is a connection, not a rebuild.
How does this approach map to my R&D workflow?
It maps directly onto stage-gate work. A prior art agent returns grounded references, a white space analysis returns a defensible read of unclaimed territory, a freedom-to-operate query runs against live patent data, and regulatory tracking stays current through the vertical layer. Each workflow is answered by the domain engine rather than the frozen general model.
If everyone can access the same AI models, where is the competitive advantage?
The advantage is no longer the model, which is exactly why it is a commodity. It comes from what you connect the model to. An organization that wires its general AI to a domain-oriented R&D intelligence layer gets grounded, defensible answers, while one relying on the model alone gets fluent guesses.

For most of the past three decades, the corporate IP team occupied a clear position near the end of the innovation process. Research and development explored a concept, leadership committed resources, scientists and engineers built the product, and only then did the work reach IP for protection, prosecution, and portfolio management. IP was a service function, expert and essential, but downstream of the decisions that mattered most. That sequence has quietly inverted. Today R&D comes to IP before resources are committed, asking what already exists in the patent record and treating the answer as a go or no-go signal on whether to pursue an idea at all. A prior art search is no longer just a legal precaution. It has become a strategic input that shapes which programs get funded, which get redirected, and which get killed before a dollar is spent.
This is a meaningful elevation of the IP team's role, and in most organizations it happened by default rather than by design. The mandate expanded because R&D became too expensive and too risky to pursue on instinct. The data and the tooling underneath the IP function, however, did not expand with it. The team is now being asked forward-looking strategic questions and is answering them with the one dataset it has always owned: the patent record. That mismatch between the question being asked and the data available to answer it is the source of a specific, costly, and underappreciated error. It has a name worth retiring from strategic vocabulary: the white space fallacy, the assumption that an empty region of the patent map is an open opportunity.
The stakes are higher than the tooling reflects
The reason this matters is that the decisions riding on these analyses are enormous, and the base rates for innovation are unforgiving. Failure rates across corporate R&D are persistently high. Industry research has long pegged new product failure somewhere between a third and half of all launches, and a substantial share of R&D projects never reach production at all. These failures have many causes, but a recurring and underexamined one is the practice of validating technical opportunity through patent analysis while leaving commercial opportunity unvalidated. A program clears the patent landscape, looks open, and proceeds, only to discover that the space was empty for reasons the patent record never showed. When the IP team's answer is steering investment direction, the cost of an incomplete map is no longer a missed filing. It is a misallocated research budget and a multi-year bet placed in the wrong direction.
White space and opportunity space are not the same thing
The cleanest way to see the error is to picture two overlapping circles. The first is patent white space, the regions of a technology landscape where few or no active patents exist. The second is commercial opportunity, the areas where genuine market demand and commercial momentum are forming. The portfolio every organization actually wants sits in the overlap, where a defensible technical position meets real commercial pull. That overlap is a narrow slice, and most teams cannot see it clearly because they are looking at only one of the two circles.
The reason patent white space gets mistaken for opportunity is structural rather than careless. Patent data is the dataset the IP team owns, the tool it has on hand, and the answer it can produce on demand. So the strategic question silently narrows from where should we invest to where is the patent map empty, and those two questions only sometimes have the same answer. The narrowing is invisible because it happens inside the framing of the analysis, not in its conclusions. Everyone in the room believes they are discussing opportunity. They are actually discussing patent density.
An empty region of the patent map can mean two very different things, and distinguishing between them is the whole game. It can be open for a reason, because there is no market demand, because the underlying science does not work yet, or because the unit economics never close. Easy to patent does not mean possible to monetize, and a clear space on the map can simply be a place no one has bothered to claim because there is nothing there worth claiming. Alternatively, the empty space can be a trap of the opposite kind, a region where competitors are very much active but moving through channels that never touch the patent system: trade secrets, defensive publications, or simply faster commercial execution that outruns the filing timeline. In both cases the patent map looks identical. It looks open. Only data drawn from outside the patent system can tell you which kind of empty you are actually looking at, and the two demand completely different strategic responses.
The inverse error is just as expensive and far less discussed. Some of the most contested, patent-dense regions of a landscape are exactly where the market is moving, and exactly where a given organization may be dangerously under-protected. A crowded patent map instinctively reads as a closed door, a market already won by incumbents. But density is a measure of competitive intensity, not of whether the opportunity is worth pursuing. Some of the most commercially urgent positions a company can take are in crowded spaces where the organization holds a real technical advantage but has under-filed relative to the competition. Reading crowdedness as a stop sign can forfeit exactly the positions most worth fighting for.
A patent is a twenty-year bet placed with rear-view data
Underneath the white space problem sits a deeper structural mismatch, this one about time. A patent is a roughly twenty-year commitment. That makes it one of the most forward-looking instruments a company holds, a claim staked on what will matter for two decades. Yet the patent record itself is one of the most backward-looking datasets available to anyone. Applications publish around eighteen months after they are filed, and the decisions behind them were made well before that. By the time a filing is visible in the public record, it describes a strategic choice that may be two or three years old. Patents are lagging indicators, sometimes by years, as applications crawl through prosecution. A team that validates a long-horizon investment using only existing patents is steering a twenty-year bet with a dataset that describes where the field was, not where it is going.
The question the IP team is increasingly asked to answer is whether a given portfolio or technology area will still matter in five to ten years. Answering that honestly requires three categories of signal that the patent record either omits entirely or reports too late to be useful.
The first is scientific momentum. Peer-reviewed papers, preprints, grant awards, and clinical activity reveal where the underlying technology is heading long before any of it reaches a patent application. Preprints in particular can surface a competitor's technical direction months to years ahead of the corresponding filing, because the science is published when it is done, not when the legal strategy is finalized. A field rich in recent publication but thin on filings is frequently an emerging opportunity, an early window in which an organization can establish a position before the patent landscape fills in and the easy ground is taken. To a patent-only view, that same field registers as white space and risks being dismissed as empty, when it is in fact the most valuable kind of crowded: crowded with science, not yet with claims.
The second is commercial signal. Venture funding, startup formation, mergers and acquisitions, corporate disclosures, and product launches reveal where commercial conviction is forming, frequently well ahead of patent activity. A technology domain showing minimal patent filings but hundreds of millions of dollars in aggregate venture funding is not white space. It is a market building momentum through channels that patent analytics simply cannot see. When an acquirer buys a startup, the strategic implication for every competitor in the space is immediate, but the patent assignment record may take months to update, and the commercial rationale for the deal, which market is being targeted, which product lines will expand, which competing approaches are being consolidated, never enters the patent data at all. That intelligence lives in deal records, regulatory filings, and corporate disclosures, in a layer of the landscape the patent-only team never sees.
The third is forward indicators, the signals that point at intent before it materializes as anything protectable. Regulatory filings, clinical pipelines, market intelligence, and hiring patterns all belong here. Hiring is among the most underused signals of all. The engineering and research roles a company is staffing frequently describe, in the job specifications themselves, exactly what the organization is building, and they appear long before any of that work surfaces as a filing. A competitor assembling a team around a specific technical capability is making a far earlier and often far clearer statement of direction than anything that will eventually reach a patent office.
None of this argues for abandoning patent data. Global patents remain the foundation, the authoritative record of what has actually been claimed and protected, and no serious analysis proceeds without them. The argument is narrower and harder to dismiss: patents are necessary but not sufficient for the strategic questions IP teams are now expected to answer. The foundation is solid. The problem is that three of the four walls are missing, and the team is being asked to assess the whole structure from the foundation alone.
Why the gap persists when it is so clearly understood
If the gap is this obvious, the fair question is why it endures across so many sophisticated organizations. The answer is mostly structural, not a failure of intelligence or diligence. Patent data is, for the typical IP team, the only native dataset it owns. It arrives through tools built for patent prosecution and portfolio management, instruments designed for IP attorneys running episodic, filing-driven workflows. Those tools are genuinely excellent at the job they were built to do. They were simply never built to answer strategic, forward-looking, commercially grounded questions, because those questions were not part of the IP team's mandate when the tools were designed.
The result is a quiet optimization toward the measurable. Teams optimize for the data they can see, and white space becomes the proxy for opportunity precisely because white space is the one thing the available tooling can actually measure. Scientific momentum, commercial conviction, and forward intent are harder to see not because they are less important but because they live in datasets the IP team's tools were never wired to ingest. The gap persists because closing it has historically meant stitching together multiple disconnected platforms by hand, a manual integration burden that most teams cannot sustain quarter after quarter. So the easier path wins, and the patent map stands in for the opportunity map by default.
Closing the gap, then, is not a matter of working harder inside the patent record. No amount of additional rigor applied to a patent-only dataset produces the signals that dataset does not contain. The fix is to put the other datasets on the same surface as the patent data, so that both circles can finally be examined together rather than one at a time, and so the overlap, the actual opportunity space, becomes visible rather than inferred.
Where this is heading
The platforms built for this problem treat patents, scientific literature, and commercial signals not as separate vendor silos to be reconciled by analysts but as a single intelligence substrate. Cypris was built specifically for this, an enterprise R&D intelligence platform that unifies more than 500 million patents and scientific papers alongside commercial and market signals, grounded in a proprietary R&D ontology and serving hundreds of enterprise customers and thousands of R&D and IP professionals across Fortune 500 companies. The application most relevant to the white space problem is exactly the overlap: surfacing the gaps between heavy patent activity and heavy publication activity, and the spaces where academic or commercial momentum is building but filings have not yet appeared. Those patterns are the opportunity space, and they are invisible inside any single-source tool by construction, because no single source contains both halves of the picture.
The more recent shift is from periodic analysis toward continuous intelligence. In June 2026 Cypris launched Agentic Monitoring, which runs continuously across patent offices, scientific literature, regulatory bodies, mergers and acquisitions, product launches, grant awards, and corporate news, delivering filtered and contextualized intelligence on a defined cadence rather than waiting for a quarterly manual rebuild. The significance is not the automation in itself. It is that the strategic questions reaching the IP team do not pause between reporting cycles. Competitors hire, raise, publish, and acquire continuously, and an intelligence model that refreshes once a quarter is structurally behind the landscape it is meant to describe. Continuous monitoring closes the timing gap on the same logic that integrated data closes the coverage gap.
The role of the corporate IP team has evolved into something genuinely strategic. The mandate, the data, and the tooling are only now beginning to catch up to it. The organizations that close that gap first will be the ones making forward decisions with a forward-looking map, while their competitors are still reading the rear-view mirror and calling it the road ahead.
FAQ
What is the difference between patent white space and commercial opportunity space?
Patent white space refers to regions of a technology landscape where few or no active patents exist. Commercial opportunity space refers to areas where genuine market demand and commercial momentum are forming. The two overlap only partially, and the highest-value IP portfolios sit in the intersection where a defensible technical position meets real commercial demand. Patent data alone cannot identify that intersection because it captures only one of the two dimensions, which is why empty patent regions are routinely mistaken for open opportunities.
What is the white space fallacy?
The white space fallacy is the assumption that an empty region of the patent map represents an open commercial opportunity. An absence of patents is a starting point for investigation, not a validated opportunity. A space can be empty because there is no market, because the underlying science does not yet work, or because competitors are operating outside the patent system through trade secrets, defensive publications, or faster commercial execution. Patent data cannot distinguish between these cases, and each one demands a completely different strategic response.
Why can patent data not answer strategic R&D questions on its own?
A patent is a roughly twenty-year commitment, which makes it a forward-looking instrument, while the patent record is a backward-looking dataset that publishes filings about eighteen months after submission and reflects decisions made earlier still. Patents are lagging indicators, sometimes by years. Answering whether a technology area will still matter in five to ten years requires scientific momentum, commercial signals, and forward indicators that the patent record either omits entirely or reports too late to act on.
Has the role of the corporate IP team actually changed?
Yes, and substantially. The IP team historically protected innovations after R&D produced them, sitting downstream of the decisions that mattered. Increasingly, R&D consults IP before committing resources and treats the resulting landscape analysis as a strategic go or no-go signal. The IP function has become a strategic decision input that shapes investment direction, even though the underlying data and tooling were originally built for patent prosecution and portfolio management rather than strategy.
What datasets do IP teams need beyond patents?
Three categories. Scientific literature, including papers, preprints, grants, and clinical activity, shows where technology is heading before filings appear. Commercial signals, including venture funding, startup formation, mergers and acquisitions, and product launches, show where commercial conviction is forming. Forward indicators, including regulatory filings, clinical pipelines, market intelligence, and hiring patterns, signal intent before it becomes protected IP. Patents remain the foundation, but these three categories supply the walls the foundation alone cannot.
Why does a field with many publications but few patents matter?
A technology area with extensive recent scientific publication but limited patent filings often represents an emerging opportunity, an early window in which an organization can establish an IP position before the landscape fills in. A patent-only view registers this same area as white space and may dismiss it as empty, missing the signal entirely. The space is not empty. It is crowded with science that has not yet converted into claims.
Can hiring patterns really indicate competitive activity?
Yes, and they are among the earliest signals available. The engineering and research roles a company staffs frequently describe, in the job specifications themselves, exactly what the company is building. Because hiring precedes filing by a considerable margin, a competitor's hiring activity can reveal technical direction months or years before any of that work surfaces in the patent record.
Why does a crowded patent area still matter strategically?
A patent-dense area instinctively reads as a closed market, but contested areas are often exactly where the market is moving and where an organization may be under-protected. Density signals competitive intensity, not the absence of opportunity. Treating a crowded map as a closed door can forfeit positions where a company holds a real technical advantage but has under-filed, which can be as costly an error as treating an empty map as an open opportunity.
Why does this gap persist if it is so well understood?
The gap is structural rather than a failure of judgment. Patent data is the only native dataset most IP teams own, accessed through tools built for prosecution and portfolio management. Teams optimize for the data they can see, so white space becomes a proxy for opportunity because it is the dimension the available tooling can actually measure. Historically, closing the gap meant manually stitching together disconnected platforms quarter after quarter, a burden most teams could not sustain, so the patent-only default persisted.
How are platforms addressing the patent-only limitation?
Purpose-built R&D intelligence platforms unify patents, scientific literature, and commercial signals into a single searchable substrate rather than separate tools requiring manual reconciliation. This allows teams to see the overlap between technical defensibility and commercial momentum directly rather than inferring it. The emerging direction is continuous monitoring across patents, literature, regulatory activity, mergers and acquisitions, and corporate news, replacing periodic manual analysis with always-on intelligence that keeps pace with a landscape that never stops moving.

Small modular reactors have moved from a policy talking point to a genuine industrial race, and their patent landscape is distinctive because "modular" is as much a manufacturing and business-model claim as it is a reactor-physics one. A small modular reactor is conventionally defined as a nuclear fission unit rated at or below roughly 300 MWe and engineered for factory fabrication and modular deployment, with microreactors forming a further, smaller subcategory typically at or below about 20 MWe¹,². The intellectual property divides across several regions, each a distinct area of patenting: reactor core and fuel design, spanning light-water designs and advanced non-light-water designs using gas, liquid metal, or molten salt as a coolant; passive safety systems, which rely on natural circulation, integral primary-system design, and large coolant inventory per unit of power rather than powered pumps and operator action — a design philosophy explicitly framed in the literature as a direct lesson from prior operating experience³,⁴; factory fabrication and modular-construction methods, the core cost and schedule thesis behind SMRs; and grid, thermal-storage, and data-center integration. Because a deployable SMR project depends on all of these layers, and because reactor types differ fundamentally in coolant and fuel choice, freedom-to-operate and white space analysis must span reactor type and layer together.
Global deployment status is best read from primary trackers such as the IAEA's Advanced Reactors Information System and coordinated European Commission Joint Research Centre analysis, which draws directly on that database to map the SMR ecosystem, rather than from market-research aggregation⁵. Reliable, precise, primary-sourced counts of reactors currently operating, under construction, or in licensing were not confirmed against an authoritative tracker in this research pass, so specific status figures should be verified against ARIS or the equivalent national regulator's own docket before being cited as current. On the fuel side, HALEU (high-assay low-enriched uranium, enriched to roughly 5–20% U-235) is a recognized supply-chain bottleneck for most advanced non-light-water designs; a European Parliament briefing, citing the US program, reports that Centrus Energy produced the first US HALEU in over 70 years under the Department of Energy's HALEU Availability Program, targeting roughly 900 kg per year toward 2030, though this is a secondary (EU) rendering of the US disclosure rather than the DOE's own primary document⁶. Because applications publish about eighteen months after filing, the most recent passive-safety and modular-fabrication filings are under-represented, so the current frontier is more active than granted-patent counts suggest.
The dominant driver of near-term commercial interest is electricity demand from AI data centers, and the clearest primary-sourced example is Google's own announcement of what it described as the first corporate agreement to purchase nuclear energy from multiple SMRs, an order for up to 500 MW of capacity from Kairos Power with a first unit targeted around 2030 — a target date, not a regulator-confirmed operating date⁷,⁸. Other widely cited data-center nuclear commitments from additional technology companies were sourced in this pass from secondary reporting rather than each company's own press release or the relevant utility's regulatory filing, and should be confirmed against those primary sources before being treated as settled. On the economics side, peer-reviewed work provides the methodological backbone for SMR cost analysis, and the recurring finding is that modularity and factory learning are the central economic lever behind SMR cost claims but remain empirically unproven, since no SMR has yet actually been built at commercial scale to validate the factory-learning thesis⁹.
The strategic picture turns on which reactor type and which layer is hardest to design around, and here the patent corpus itself requires a significant caveat. A patent search on the literal term "SMR" is heavily contaminated by an entirely unrelated field that shares the same abbreviation — steam methane reforming, a chemical-reactor process — such that a raw, unfiltered ranking is dominated by petrochemical entities that are not nuclear SMR filers at all. Once filtered to clearly nuclear assignees, the genuine SMR patent landscape is led by reactor developers such as Westinghouse, NuScale, TerraPower, and BWXT, alongside academic and national-institution filers working on molten-salt designs. Passive safety-system IP is widely regarded as the single most technically intensive and contested domain in SMR development, since it is central to both regulatory approval and the reduced-footprint site design that makes SMRs viable near data centers and other non-traditional locations. Beyond safety systems, the white space includes non-light-water reactor types that remain earlier in development and less crowded than light-water SMRs; HALEU fuel supply chain and fabrication IP, a genuine bottleneck across nearly every advanced design; and the thermal and electrical integration systems that couple a reactor to a data center's variable, high-density cooling and power loads. Reading the landscape by reactor type, layer, and owner — after filtering out the steam-methane-reforming noise — and tracking both the patents and the underlying nuclear-engineering research, is what separates a workable deployment position from a blocked one.
Where the small modular reactor white space is
Non-light-water reactor types. Gas-cooled, liquid-metal-cooled, and molten-salt SMR designs are earlier in development and less crowded than light-water designs, offering higher-temperature output and, in some cases, simplified passive safety.
HALEU fuel supply chain and fabrication. High-assay low-enriched uranium fuel is required by most advanced non-light-water designs and remains a genuine, primary-sourced supply-chain bottleneck, making fuel-fabrication and enrichment IP a distinct, high-value layer⁶.
Factory fabrication and modular construction. Methods that close the cost and schedule gap between a first-of-a-kind unit and Nth-of-a-kind serial production are the core economic thesis of SMRs, and peer-reviewed economics literature confirms this thesis remains empirically unvalidated at scale — a genuine open question, not settled fact⁹.
Data-center thermal and electrical integration. Coupling reactor heat-rejection and power output to a data center's variable, high-density cooling and compute loads is an emerging, largely unclaimed layer distinct from conventional grid integration.
Verified project-status tracking. Because primary-sourced operating/under-construction/licensing status is scarce relative to the volume of announcements, and because the patent corpus itself requires filtering against an unrelated identically-named chemical process, a rigorously verified view of the field is itself a differentiator.
How AI-powered landscape and white space analysis helps
Resolving a landscape that spans multiple reactor coolant types, passive safety-system engineering, fuel supply chain, and an entirely new data-center integration layer — while filtering out an unrelated, identically-abbreviated chemical-process field — requires more than keyword search. AI-powered analysis addresses this with semantic search that clusters activity by reactor type and layer while distinguishing nuclear SMR filings from steam-methane-reforming noise, attribution that normalizes reactor-developer, utility, and technology-company filers to canonical entities, and continuous monitoring that keeps pace with a field where licensing milestones and data-center power deals are both moving quickly. Because nuclear-engineering advances appear in scientific and regulatory literature before they translate into patents, reading both patents and literature gives the earliest signal of which reactor type and layer is actually closing the gap to commercial deployment.
The competitive landscape by the numbers
The raw "SMR" patent corpus is dominated by steam methane reforming and general chemical-reactor art rather than nuclear small modular reactors — the top unfiltered assignees include major petrochemical and catalysis companies that have no connection to nuclear technology, and this contamination means raw top-N assignee or geography rankings from an unfiltered query should not be presented as a nuclear SMR landscape (Cypris corpus, indicative; 2025–26 partial). Filtering to clearly nuclear-specific assignees surfaces the genuine SMR reactor-developer landscape led by Westinghouse, NuScale, TerraPower, and BWXT (developer of the mPower design), with academic and national-institution filers active in molten-salt-specific IP (Cypris corpus, indicative; 2025–26 partial). Geography in the filtered nuclear-specific set skews toward China and the United States. A reliable coolant-type and layer-specific split, and a clean total patent-family count, could not be produced from the contaminated raw corpus in this pass and are not presented here as authoritative; a follow-on query built on a nuclear-specific classification filter (rather than the "SMR" keyword alone) is needed to produce a trustworthy count.
Where Cypris fits
Cypris runs patent landscape and white space analysis for fast-deploying energy fields such as small modular reactors across a corpus of more than 500 million patents and scientific papers, organized through a proprietary R&D ontology. The ontology clusters activity by reactor type, light-water, gas-cooled, liquid-metal-cooled, and molten-salt, and by layer, core and fuel design, passive safety, factory fabrication, and grid/data-center integration, and normalizes reactor-developer, utility, and technology-company filers to canonical entities — critically, distinguishing genuine nuclear SMR filings from the unrelated steam-methane-reforming field that shares the same abbreviation — so a team can resolve which reactor types and layers are crowded and which remain open as white space. Semantic search across patents and scientific literature connects filings to the underlying nuclear-engineering research, which is where SMR advances appear first. Cypris Q, the platform's agentic layer, lets teams run landscape and white space analysis conversationally and chain the clustering, attribution, and gap analysis, and Agentic Monitoring tracks a defined reactor type or layer over time and flags new patents and papers as they publish. Cypris provides enterprise API partnerships with OpenAI, Anthropic, and Google, and is built with enterprise-grade security. Cypris serves hundreds of enterprise customers across pharmaceuticals, chemicals, advanced materials, energy, and other regulated industries.
FAQ
What is a small modular reactor? A small modular reactor is a nuclear fission unit generally rated at or below roughly 300 MWe (microreactors at or below about 20 MWe), engineered so its major components can be built on an assembly line in a factory and shipped to site rather than constructed piece by piece in place¹,². This factory-first approach is the core cost and schedule thesis behind SMRs, though peer-reviewed economics literature notes it remains empirically unvalidated since no SMR has yet been built at commercial scale⁹. Designs span both light-water and advanced non-light-water coolant types.
Why are small modular reactors a patenting hotspot now? Small modular reactors are a patenting hotspot now because AI data centers' rapidly growing electricity demand has made carbon-free, co-locatable baseload power commercially urgent — Google's own announcement of an order for up to 500 MW from Kairos Power is the clearest primary-sourced example of this trend⁷,⁸. That deployment pressure is driving filings across safety systems, fuel design, and data-center integration.
What layers does the SMR patent landscape cover? The landscape covers reactor core and fuel design, passive safety systems, factory fabrication and modular-construction methods, and grid and data-center integration. Each is a distinct region of patenting held by different developers, utilities, and technology companies. Freedom-to-operate and white space analysis must span reactor type and layer together.
Why is the "SMR" patent corpus hard to search accurately? The "SMR" patent corpus is hard to search accurately because the abbreviation is shared with steam methane reforming, an unrelated chemical process for producing hydrogen, and a raw keyword search returns a corpus dominated by petrochemical and catalysis companies rather than nuclear reactor developers. Filtering to nuclear-specific classification and assignees is required to see the genuine small modular reactor landscape, led by developers such as Westinghouse, NuScale, TerraPower, and BWXT. This is a significant, easy-to-miss data-quality issue in SMR patent analysis.
Why are passive safety systems the most contested layer? Passive safety systems are the most contested layer because they rely on natural circulation and integral primary-system design rather than powered pumps and operator action, a design philosophy explicitly developed as a lesson from prior operating experience³,⁴, and because it is central both to regulatory approval and to the reduced-footprint site design that makes SMRs viable in non-traditional locations such as data-center campuses. It is accordingly one of the most technically intensive and IP-contested domains in SMR development.
Where is the white space in small modular reactors? The white space includes non-light-water reactor types, HALEU fuel supply chain and fabrication, factory fabrication and modular construction methods (an economically unproven thesis worth backing with real data), data-center thermal and electrical integration, and rigorously verified project-status tracking. Light-water SMR designs and core passive-safety concepts are comparatively more developed. The newer reactor types and the data-center integration layer are the most open ground.
Why is HALEU fuel a bottleneck? HALEU, or high-assay low-enriched uranium, is required by most advanced non-light-water SMR designs, and while the US has begun domestic production under the DOE's HALEU Availability Program, reported capacity remains modest (on the order of 900 kg per year targeted toward 2030) relative to the number of designs that depend on it⁶. This makes fuel supply chain and fabrication IP a distinct, high-value layer independent of reactor design itself.
Why does SMR analysis need scientific literature? SMR analysis needs scientific literature because reactor-physics, fuel, and passive-safety advances appear in nuclear-engineering research and regulatory technical literature before they are patented, and because much of the public narrative around SMR deployment status and data-center deals outruns what is confirmed in primary regulatory or company disclosures. Analyzing patents alone gives a lagging view. Cypris analyzes both across more than 500 million patents and scientific papers.
Which teams use small modular reactor patent landscape analysis? Small modular reactor patent landscape analysis is used by R&D, IP, and strategy teams at reactor developers, utilities, and data-center and technology companies exploring co-located nuclear power, as well as investors and policymakers. Because the landscape spans multiple reactor types at different licensing and deployment stages, and because raw keyword search is contaminated by an unrelated chemical-process field, structured analysis is essential. Cypris serves hundreds of enterprise customers across energy and other research-intensive industries.
Endnotes
- Friedman E. Small Modular Reactors (SMRs). Oxford University Press eBooks. DOI: 10.1093/9780198925811.003.0023.
- Sinha V. Small Modular Reactors (SMRs) and Microreactors: Understanding the Major Designs Shaping the Future of Nuclear Energy. Zenodo. DOI: 10.5281/zenodo.20710566.
- Ingersoll DT. Passive Safety Features for Small Modular Reactors. World Scientific eBooks. DOI: 10.1142/9789814365932_0012.
- Ilyas M, Aydoğan F, Butt HN, Ahmad M. Assessment of passive safety system of a Small Modular Reactor (SMR). Annals of Nuclear Energy. DOI: 10.1016/j.anucene.2016.07.018.
- European Commission Joint Research Centre. An exploratory analysis of the Small Modular Reactor ecosystem (drawing on IAEA ARIS). publications.jrc.ec.europa.eu.
- European Parliament Research Service (EPRS). Strategic autonomy and the future of nuclear energy in the EU, citing the US DOE HALEU Availability Program and Centrus Energy production. europarl.europa.eu.
- Google. Google signs advanced nuclear clean energy agreement with Kairos Power. Company blog announcement, October 2024.
- Kairos Power. Google and Kairos Power Partner to Deploy 500 MW of Clean Electricity Generation. Company press release.
- Locatelli G, Mignacca B. Economics and finance of Small Modular Reactors: A systematic review and research agenda. Renewable and Sustainable Energy Reviews. DOI: 10.1016/j.rser.2019.109519.
- Cypris platform corpus analysis, small modular reactor patent families (nuclear-filtered where noted). Indicative figures; 2025–2026 partial.

Sustainable aviation fuel has moved from pilot projects to a mandated market, and its patent landscape is distinctive because SAF is not a single technology but a set of competing production routes, each with its own feedstocks, catalysts, and process chemistry. Peer-reviewed technical reviews lay out the route taxonomy: the hydroprocessed-ester-and-fatty-acid route converts waste oils and fats into jet fuel and is currently the most mature; the Fischer-Tropsch route gasifies biomass or waste into synthesis gas and rebuilds it into hydrocarbons; the alcohol-to-jet route converts ethanol or other alcohols into jet-range molecules; and the synthetic power-to-liquid route, including methanol-mediated pathways, combines captured carbon dioxide with green hydrogen to make e-fuels with no biological feedstock at all.¹,²,³,⁴,⁷ Because each route is a distinct region of patenting, freedom-to-operate and white space analysis must treat SAF as several landscapes at once, spanning feedstock pretreatment, catalysts, conversion processes, and upgrading.
The landscape is being pulled forward by regulation more directly than most. Under the European Union's ReFuelEU Aviation regulation, the sustainable share of aviation fuel supplied at EU airports rises stepwise to 70 percent by 2050, with a dedicated sub-obligation for synthetic e-fuels and an anti-tankering rule requiring airlines to uplift most of their fuel where they operate; Switzerland adopted the ReFuelEU framework from January 1, 2026.⁹ This creates both a deadline and a guaranteed market against a very large baseline, since global commercial jet-fuel demand is on the order of 100 billion gallons a year and is projected to rise substantially by 2050.¹ The near-term response has concentrated in the waste-oil route because it is the most mature,⁴,⁵ but the mandates specifically favor synthetic e-fuels in the longer term, which is steering research and filings toward the power-to-liquid route and its underlying carbon-conversion and catalysis challenges. The patent record shows this tension clearly: across the Cypris corpus of more than 500 million patents and scientific papers, the SAF space holds roughly 6,000 de-duplicated families and grew about 3.6 times between 2022 and 2024, and on an indicative basis the Fischer-Tropsch and e-fuel routes lead patent activity, ahead of hydroprocessed waste oils, with alcohol-to-jet the smallest slice, even though the waste-oil route currently leads in deployed production capacity, a divergence between where filing and where building are concentrated. The most active assignees span engine makers, refining-and-catalysis licensors, and route pure-plays, and the United States leads on geography, followed by the United Kingdom, China, France, and the Nordic producers. Because applications publish about eighteen months after filing, the most recent catalyst and e-fuel filings are under-represented (2025 and 2026 counts are partial), so the current frontier is more active than granted-patent counts suggest.
The strategic question is which route and layer to back, and the white space sits where cost and feedstock constraints are hardest. The waste-oil route is limited by feedstock availability, so its white space is narrower; the Fischer-Tropsch and alcohol-to-jet routes turn on catalyst performance and process integration;²,³,⁸ and the synthetic e-fuel route, though earliest and most expensive, is the one the mandates most favor and the one with the most open, high-value IP, particularly in the catalysts and process designs that lower the cost of converting carbon dioxide and hydrogen into jet fuel.⁶,⁷ Reading the landscape by route, feedstock, catalyst, and process, and tracking both the patents and the underlying chemistry research, is what separates a crowded region from an open one.
Where the SAF white space is
Synthetic e-fuel catalysis. Catalysts and process designs that lower the cost of converting captured carbon dioxide and green hydrogen into jet-range hydrocarbons are the most favored by mandate and among the most open, high-value targets.⁶,⁷
Alcohol-to-jet conversion. Improved catalysts and process integration for converting alcohols to jet-range molecules are an active, still-developing route.⁸
Fischer-Tropsch from waste and biomass. Gasification, syngas conditioning, and Fischer-Tropsch catalysis for waste and biomass feedstocks are a distinct, contested layer.²,³
Feedstock flexibility and pretreatment. Technologies that broaden or pretreat feedstocks, easing the supply constraint on mature routes, are a differentiated area.⁴
Process intensification and integration. Designs that integrate steps, cut energy use, and lower capital cost are where scale-up economics are decided.⁵
How AI-powered landscape and white space analysis helps
Resolving a landscape that spans several production routes, each with its own feedstocks, catalysts, and processes, under a moving regulatory timeline, requires more than keyword search. AI-powered analysis addresses this with semantic search that clusters activity by route, feedstock, catalyst, and process across varied terminology, attribution that normalizes filers to canonical entities and tracks new entrants, and continuous monitoring that keeps pace with a mandate-driven surge. Because SAF advances appear in scientific and catalysis literature before they are patented, reading both patents and literature gives the earliest signal of where scalable routes are emerging.
Where Cypris fits
Cypris runs patent landscape and white space analysis for multi-route fields such as sustainable aviation fuel across a corpus of more than 500 million patents and scientific papers, organized through a proprietary R&D ontology. The ontology clusters activity by production route, waste-oil, Fischer-Tropsch, alcohol-to-jet, and synthetic e-fuel, and by layer, feedstock, catalyst, conversion, and upgrading, and normalizes filers to canonical entities, so a team can resolve which routes and layers are crowded and which remain open as white space, and can track new entrants as the field scales. Semantic search across patents and scientific literature connects filings to the underlying catalysis and process research, which is where SAF advances appear first. Cypris Q, the platform's agentic layer, lets teams run landscape and white space analysis conversationally and chain the clustering, attribution, and gap analysis, and Agentic Monitoring tracks a defined route over time and flags new patents and papers as they publish. Cypris provides enterprise API partnerships with OpenAI, Anthropic, and Google, and is built with enterprise-grade security. Cypris serves hundreds of enterprise customers across pharmaceuticals, chemicals, advanced materials, energy, and other regulated industries.
FAQ
What is the sustainable aviation fuel patent landscape? The sustainable aviation fuel patent landscape is the set of patents covering the several routes used to make jet fuel with lower lifecycle emissions, including hydroprocessed waste oils, Fischer-Tropsch fuels, alcohol-to-jet, and synthetic power-to-liquid e-fuels. Each route has distinct feedstocks, catalysts, and processes. It is best understood as several landscapes rather than one.
Why is regulation shaping SAF patenting? Regulation shapes SAF patenting because binding blending mandates require a rising share of sustainable aviation fuel over the coming decades, reaching 70 percent by 2050 under the EU ReFuelEU Aviation regulation, with a dedicated sub-mandate for synthetic e-fuels. Near-term activity has concentrated in the mature waste-oil route, while the mandates steer longer-term research toward e-fuels. The patent record tracks this policy pull closely.
What production routes does the SAF landscape cover? The SAF landscape covers hydroprocessed waste oils and fats, Fischer-Tropsch fuels from gasified biomass or waste, alcohol-to-jet conversion, and synthetic power-to-liquid e-fuels made from captured carbon dioxide and green hydrogen. Each is a distinct region of patenting. Freedom-to-operate and white space analysis must treat them separately.
Where is the white space in SAF? The white space sits in synthetic e-fuel catalysis, alcohol-to-jet conversion, Fischer-Tropsch from waste and biomass, feedstock flexibility and pretreatment, and process intensification. The mature waste-oil route is comparatively crowded and feedstock-limited. The most open, high-value opportunities are in the e-fuel catalysts and processes the mandates most favor.
Why is the synthetic e-fuel route strategically important? The synthetic e-fuel route is strategically important because the mandates specifically favor it in the longer term, it has no biological feedstock limit, and it is the least mature and most expensive route, which leaves the most open, high-value IP. The central challenge is lowering the cost of converting carbon dioxide and hydrogen into jet fuel. That is where much of the defensible catalysis and process IP is concentrating.
Why does the patent record differ from deployed capacity in SAF? The patent record differs from deployed capacity because filing tends to run ahead of building. In the Cypris corpus the Fischer-Tropsch and e-fuel routes lead in patent activity, even though the hydroprocessed waste-oil route currently leads in installed production capacity. That divergence signals where developers expect the next phase of growth.
What software helps analyze the sustainable aviation fuel patent landscape? Software for the SAF landscape should cluster activity by production route and process layer, resolve filers to canonical owners, search patents and scientific literature semantically, and monitor a mandate-driven field continuously. Cypris does this across more than 500 million patents and scientific papers using a proprietary R&D ontology, semantic search, Cypris Q, and Agentic Monitoring.
Which teams use SAF patent landscape analysis? SAF patent landscape analysis is used by R&D, innovation, IP, and strategy teams at fuel producers, chemicals and catalysis companies, airlines and energy majors, and their partners, as well as investors and policymakers. It informs which route to back, where to file, and where competitors are concentrated. Cypris serves hundreds of enterprise customers across chemicals, energy, advanced materials, and other regulated industries.
Endnotes
- Heyne, J., Holladay, J., & Abdullah, Z. (2020). Sustainable aviation fuel: review of technical pathways. Pacific Northwest National Laboratory / U.S. Department of Energy, Bioenergy Technologies Office. https://doi.org/10.2172/1660415
- Zhang, X., Zheng, Y., Li, J., & Wang, X. (2025). Research advances and future perspectives in Fischer-Tropsch synthesis for sustainable aviation fuel. Sustainable Energy & Fuels. https://doi.org/10.1039/d5se01412c
- Vreugdenhil, B., Boymans, E., Viar, H., et al. (2025). Syngas to sustainable aviation fuel: emerging catalysts and routes. Applied Catalysis A: General. https://doi.org/10.1016/j.apcata.2025.120554
- Chang, K., Ng, J., Japar, W. M. A. W., et al. (2026). Lipid feedstocks for sustainable aviation fuel via HEFA: status and challenges. Renewable and Sustainable Energy Reviews. https://doi.org/10.1016/j.rser.2026.117006
- Gómez, J., & Gyandoh, D. (2025). Techno-economic analysis of HEFA and lignocellulosic biomass conversion for sustainable aviation fuel. Applied Energy. https://doi.org/10.1016/j.apenergy.2025.126421
- Riaz, A., Qyyum, M. A., Al-Muhtaseb, A. H., Al-Jahwari, F., & Saeed, A. (2026). Carbon-derived and biomass-based sustainable aviation fuel pathways: a comparative techno-economic and life-cycle review for aviation decarbonization. Carbon Capture Science & Technology. https://doi.org/10.1016/j.ccst.2026.100641
- Karlsruhe Institute of Technology (2025). Sustainable aviation fuel production via the methanol pathway: a technical review. Sustainable Energy & Fuels. https://doi.org/10.5445/ir/1000187428
- Probabilistic technoeconomic analysis of alcohol-to-jet sustainable aviation fuel: implications for design and decision making (2026). https://doi.org/10.1088/2977-3504/ae7801/v2/review1
- European Commission, Directorate-General for Mobility and Transport. ReFuelEU Aviation. https://transport.ec.europa.eu/transport-modes/air/environment/refueleu-aviation_en

Cellular reprogramming has become one of the most closely watched areas in longevity biotechnology, and its patent landscape is distinctive because the leading approach builds directly on an already foundational technology. Full reprogramming, using the four Yamanaka factors, resets an adult cell all the way to a pluripotent, embryonic-like state; partial or transient reprogramming instead applies a subset of those factors briefly, aiming to roll back the epigenetic state of an aged cell toward a younger profile while preserving its identity and function. In animal models, partial reprogramming has ameliorated age-associated hallmarks and, in one landmark study, restored youthful epigenetic patterns and recovered vision after optic-nerve injury, evidence that framed aging partly as a loss of epigenetic information that reprogramming can help reverse.¹,² Because a rejuvenation therapy is assembled from several independently patentable pieces, the reprogramming-factor set and its ratios, the delivery system, the inducible control mechanism, the target tissue and indication, and the tools used to measure biological age, freedom-to-operate is a multi-layer, multi-owner analysis rather than a single clearance.
The foundational layer shapes everything above it. The original induced-pluripotent-stem-cell reprogramming methods, established through the forced expression of a defined set of transcription factors, sit under a well-known foundational estate that has been broadly licensed,³ and partial-reprogramming approaches inherit questions about how far that foundation reaches. Independent work has shown that epigenetic reprogramming can unlock tissue regenerative potential, reinforcing why these methods are so contested.⁴ This academic origin is visible in the ownership record: across the Cypris corpus of more than 500 million patents and scientific papers, the most active assignees in the cellular-reprogramming and induced-pluripotency space are led by academic and translational institutions, including Kyoto University, the University of California San Diego, the University of Texas System, Memorial Sloan Kettering, and Harvard, alongside cell-therapy companies, and the corpus holds on the order of 28,700 de-duplicated families, with the United States, China, and Japan the leading jurisdictions. Layered on top are newer, fast-growing estates specific to partial and transient reprogramming, cyclic and inducible expression schemes, chemical or small-molecule reprogramming that avoids transcription factors altogether, and tissue-specific delivery. Because applications publish about eighteen months after filing, the most recent reprogramming, delivery, and control filings are under-represented, so the current frontier is more active than granted-patent counts suggest.
The landscape is a well-capitalized race, and the strategic question is which layer to own. In January 2026 the field reached a milestone when the US Food and Drug Administration cleared the first human trial of a partial epigenetic reprogramming therapy, an investigational optic-neuropathy treatment; the clearance authorizes a first-in-human study and is not itself evidence of efficacy.⁹ Across the Cypris corpus, filings in this space grew from a few hundred families per year at the start of the last decade to roughly 3,200 in 2024, with 2025 counts partial because of the publication lag. Several richly funded companies are pursuing different factor sets, delivery routes, and target tissues, and a recurring challenge is to separate genuine rejuvenation, a measured reduction in biological age, from a mere slowing of decline.⁵ The durable value increasingly sits not in the general idea of reprogramming, which rests on the contested foundation, but in the specific, well-supported improvements: safe and controllable expression systems that avoid tumor risk, factor combinations and chemical alternatives, tissue-targeted delivery, and the validated biomarkers, including epigenetic clocks, used to demonstrate rejuvenation.⁶,⁷,⁸ Reading the landscape by layer and by owner, and tracking both the patents and the underlying research, is what separates a workable position from a blocked one.
What creates FTO risk in cellular reprogramming
Foundational reprogramming claims. These cover the underlying induced-pluripotency methods and factor sets, a broadly licensed foundation whose reach into partial approaches shapes everything above it.
Partial and inducible-control claims. These cover transient, cyclic, and inducible expression schemes that rejuvenate without full dedifferentiation, a fast-growing and contested layer.
Delivery claims. These cover viral vectors, lipid nanoparticles, and mRNA delivery of reprogramming factors, a distinct and separately owned layer often decisive for a therapy.
Chemical and small-molecule reprogramming claims. These cover approaches that induce rejuvenation without transcription factors, an emerging and less-crowded route.
Target, indication, and biomarker claims. These cover specific tissues and indications and the epigenetic-age measurements used to demonstrate effect, so a platform can be free for one application and blocked for another.
How AI-powered landscape and FTO analysis helps
A multi-layer, multi-owner landscape built on a contested foundation is beyond manual clearance. AI-powered analysis addresses this with semantic search that retrieves relevant foundational, partial-reprogramming, delivery, control, and target claims regardless of terminology, attribution that resolves academic and commercial owners to canonical entities and captures the license and spinout chains, claim-level analysis that separates the layers, and continuous monitoring that tracks new filings and the fast-moving research. Because reprogramming advances appear in scientific literature well before they are patented, reading both patents and literature gives the earliest warning of where the field is heading.
Where Cypris fits
Cypris runs patent landscape and freedom-to-operate analysis for multi-layer, academically rooted fields such as cellular reprogramming across a corpus of more than 500 million patents and scientific papers, organized through a proprietary R&D ontology. The ontology clusters the landscape by layer, foundational reprogramming, partial and inducible control, delivery, chemical reprogramming, and target and biomarker, and normalizes academic and commercial owners to canonical entities, so a team can trace how rights and licenses are distributed across many parties rather than read a flat list. Semantic search across patents and scientific literature surfaces relevant claims regardless of terminology and connects filings to the underlying research, which is where new factor sets, control systems, and delivery methods emerge first. Cypris Q, the platform's agentic layer, lets teams run landscape and FTO analysis conversationally and chain the attribution, clustering, and claim-level analysis across layers, and Agentic Monitoring tracks the landscape over time and flags new filings and developments as they publish. Cypris provides enterprise API partnerships with OpenAI, Anthropic, and Google, and is built with enterprise-grade security. Cypris serves hundreds of enterprise customers across pharmaceuticals, chemicals, advanced materials, energy, and other regulated industries.
FAQ
What is cellular reprogramming in the longevity context? Cellular reprogramming in the longevity context is the use of reprogramming factors to reset the epigenetic state of aged cells toward a younger profile. Partial or transient reprogramming applies a subset of the Yamanaka factors briefly, aiming to rejuvenate cells without erasing their identity. It is being pursued as an approach to age-related disease and tissue restoration.
Why is freedom-to-operate hard for reprogramming therapies? Freedom-to-operate is hard for reprogramming therapies because a therapy is assembled from several independently patentable layers, the reprogramming-factor set, the delivery system, the inducible control mechanism, the target tissue, and biomarker tools, often held by different owners on top of a foundational estate. Clearing one layer does not clear the others. FTO is therefore a multi-layer, multi-owner analysis.
How does the foundational iPSC estate affect partial reprogramming? The foundational induced-pluripotent-stem-cell estate affects partial reprogramming because partial approaches use the same reprogramming factors, so questions about how far the foundation reaches propagate into the newer methods. The foundation has been broadly licensed. Partial-reprogramming developers must consider both the foundation and the specific improvement layers.
What claim types create FTO risk in reprogramming? Five claim types create FTO risk: foundational reprogramming claims, partial and inducible-control claims, delivery claims, chemical and small-molecule reprogramming claims, and target, indication, and biomarker claims. Each covers a distinct layer and can be held by a different owner. Control systems and delivery are especially decisive.
Where is the white space in cellular reprogramming? The white space sits in safe and controllable expression systems that avoid tumor risk, chemical and small-molecule reprogramming, tissue-specific delivery, specific factor combinations, and validated biomarkers of biological age. The general concept rests on a contested foundation. The durable, defensible value is in these specific improvement and delivery layers.
Why does reprogramming analysis need scientific literature? Reprogramming analysis needs scientific literature because new factor sets, control systems, and delivery methods appear in research well before they are patented, so the literature gives the earliest signal in a fast-moving field. Analyzing patents alone gives a lagging view. Cypris analyzes both across more than 500 million patents and scientific papers.
What software helps analyze the cellular reprogramming patent landscape? Software for the cellular reprogramming landscape should resolve academic and commercial owners and license chains to canonical entities, cluster the foundational, control, delivery, and target layers, search patents and scientific literature semantically, and monitor a fast-moving field continuously. Cypris does this across more than 500 million patents and scientific papers using a proprietary R&D ontology, semantic search, Cypris Q, and Agentic Monitoring.
Which teams need reprogramming patent landscape and FTO analysis? Reprogramming patent landscape and FTO analysis is needed by R&D, IP, and business-development teams at longevity and gene-therapy companies, academic technology-transfer offices, and investors assessing rejuvenation assets. The multi-layer, contested landscape makes structured analysis essential. Cypris serves hundreds of enterprise customers across pharmaceuticals and other research-intensive industries.
This article addresses patents and freedom-to-operate and is not legal, medical, or investment advice, and contains no clinical or dosing guidance. FTO determinations should be reviewed with qualified patent counsel.
Endnotes
- Ocampo, A., Reddy, P., Izpisua Belmonte, J. C., et al. (2016). In vivo amelioration of age-associated hallmarks by partial reprogramming. Cell, 167(7). https://doi.org/10.1016/j.cell.2016.11.052
- Lu, Y., Krishnan, A., Sinclair, D. A., et al. (2020). Reprogramming to recover youthful epigenetic information and restore vision. Nature, 588. https://doi.org/10.1038/s41586-020-2975-4
- Takahashi, K., & Yamanaka, S. (2013). Induced pluripotent stem cells in medicine and biology. Development, 140(12). https://doi.org/10.1242/dev.092551
- Reddy, P., Izpisua Belmonte, J. C., & Memczak, S. (2021). Unlocking tissue regenerative potential by epigenetic reprogramming. Cell Stem Cell, 28(3). https://doi.org/10.1016/j.stem.2020.12.006
- Zhang, B., Trapp, A., Kerepesi, C., & Gladyshev, V. N. (2021). Emerging rejuvenation strategies—reducing the biological age. Aging Cell, 21(1). https://doi.org/10.1111/acel.13538
- Moqri, M., Poganik, J. R., Gladyshev, V. N., & Horvath, S. (2025). What makes biological age epigenetic clocks tick. Nature Aging. https://doi.org/10.1038/s43587-025-00833-1
- Mammalian Methylation Consortium; Horvath, S., et al. (2023). Universal DNA methylation age across mammalian tissues. Nature Aging, 3. https://doi.org/10.1038/s43587-023-00462-6
- Ferrucci, L., et al. (2019). Measuring biological aging in humans: a quest. Aging Cell, 19(2). https://doi.org/10.1111/acel.13080
- Life Biosciences (2026, January 28). Life Biosciences announces FDA clearance of IND application for ER-100 in optic neuropathies. https://www.lifebiosciences.com/life-biosciences-announces-fda-clearance-of-ind-application-for-er-100-in-optic-neuropathies

Prior art search for artificial intelligence and machine learning inventions is one of the hardest retrieval problems in patent work, for reasons specific to how AI knowledge is produced and disclosed. Prior art search establishes whether an invention is novel by finding any earlier disclosure that describes it. In most fields, the relevant disclosures are predominantly patents. In AI and machine learning, the most relevant and most recent disclosures are predominantly non-patent literature: preprints on arXiv, proceedings from conferences such as NeurIPS and ICML, open-source code and model documentation, and technical reports. These sources are published quickly and openly, often well ahead of any corresponding patent, so a prior art search confined to patent databases misses the state of the art.
The volume compounds the difficulty. AI scientific publications more than doubled from about 102,000 in 2013 to more than 242,000 in 2023, growing nearly 20 percent in the final year alone.¹ Patenting has grown even faster from a smaller base: AI patents granted worldwide rose from 3,833 in 2010 to 122,511 in 2023, an increase of almost 30 percent in the last year measured.¹ Generative AI illustrates the velocity of the literature most sharply, with related scientific publications rising from 116 in 2014 to more than 34,000 in 2023 while generative-AI patent families grew more than 800 percent over roughly the same period.² A prior art searcher in this field is therefore working against both a large and a rapidly expanding corpus, split across patent and non-patent sources.
Retrieval quality falls exactly where AI prior art needs it most. Patent retrieval is already harder than general-domain information retrieval, and controlled evaluation shows that cross-domain retrieval, finding relevant art outside the query's own technology area, performs several times worse than in-domain retrieval; one recent family-level benchmark found out-of-domain retrieval roughly five times worse than in-domain across hundreds of controlled configurations.³,⁴ AI and machine-learning methods are applied across many application domains, so relevant prior art for an AI invention is frequently located in a different field than the invention's stated use, which is precisely the cross-domain case where conventional retrieval degrades. This is the technical reason keyword and classification search alone are insufficient for AI prior art, and why dense, semantic methods have become the focus of research on patent prior art retrieval.⁵,⁶
Why AI prior art is distinctively hard
Non-patent literature dominates. The most relevant and most recent AI disclosures appear first in preprints, conference proceedings, and open-source code, so a patent-only search misses the state of the art.
Exploding volume. AI publications more than doubled to over 242,000 in 2023, and AI patents granted rose to 122,511, so the corpus a searcher must cover is both large and expanding rapidly.¹
Cross-domain dispersion. AI methods are applied across many fields, so relevant prior art is often in a different technology area than the invention, which is where retrieval degrades most.³
Fast obsolescence of terminology. AI vocabulary evolves quickly, so keyword search misses conceptually identical work described in newer or different terms.
Software-claim breadth. Algorithmic and software claims can be drafted broadly and abstractly, which makes matching a claim to its closest prior art a conceptual rather than a lexical task.
How semantic search closes the gap
Semantic search addresses each of these problems. It retrieves conceptually relevant disclosures regardless of terminology, which handles both fast-evolving vocabulary and broadly drafted software claims. Applied across both patents and scientific literature in one corpus, it covers the non-patent literature where AI prior art concentrates rather than patents alone. And because dense retrieval encodes meaning rather than surface form, it is better positioned than keyword search for the cross-domain case, retrieving relevant art from a different application area than the invention. Combined with an ontology that organizes retrieval by concept, semantic search returns a structured, high-recall view of the prior art rather than a keyword-limited sample.
Where Cypris fits
Cypris runs semantic prior art search across a corpus of more than 500 million patents and scientific papers, organized through a proprietary R&D ontology. Because the corpus spans both patents and scientific literature, Cypris covers the non-patent literature where AI and machine-learning prior art concentrates, rather than patents alone. Semantic search retrieves conceptually relevant disclosures regardless of terminology, which handles the fast-evolving vocabulary and broadly drafted software claims characteristic of AI inventions, and the ontology organizes retrieval by concept so cross-domain prior art in a different application area is surfaced rather than missed. Cypris Q, the platform's agentic layer, lets teams run and chain prior art and novelty analysis conversationally, and Agentic Monitoring tracks a technology area over time so newly published disclosures are surfaced as they appear, which matters in a field moving as fast as AI. Cypris provides enterprise API partnerships with OpenAI, Anthropic, and Google, and is built with enterprise-grade security. Cypris serves hundreds of enterprise customers across pharmaceuticals, chemicals, advanced materials, energy, and other regulated industries.
FAQ
Why is prior art search hard for AI and machine learning inventions?
Prior art search is hard for AI and machine learning inventions because the most relevant and most recent disclosures are predominantly non-patent literature, such as preprints, conference proceedings, and open-source code, which a patent-only search misses. The corpus is also large and expanding rapidly, and AI methods are dispersed across many application domains. These factors make high-recall, cross-domain retrieval essential.
Why does non-patent literature matter so much for AI prior art?
Non-patent literature matters for AI prior art because AI research is published quickly and openly, often well ahead of any corresponding patent, so the state of the art appears first in preprints, conference papers, and code. A search confined to patent databases misses these disclosures. Effective AI prior art search must cover both patents and scientific literature.
How large is the AI prior art corpus?
The AI prior art corpus is large and growing quickly. AI scientific publications more than doubled from about 102,000 in 2013 to over 242,000 in 2023, and AI patents granted worldwide rose from 3,833 in 2010 to 122,511 in 2023. Generative-AI publications alone grew from 116 in 2014 to more than 34,000 in 2023.
What makes AI prior art retrieval technically difficult?
AI prior art retrieval is technically difficult because AI methods are applied across many domains, so relevant prior art is often in a different technology area than the invention, and cross-domain retrieval performs several times worse than in-domain retrieval. One benchmark found out-of-domain retrieval roughly five times worse than in-domain. Fast-evolving terminology and broadly drafted software claims add further difficulty.
Why is keyword search insufficient for AI prior art?
Keyword search is insufficient for AI prior art because AI terminology evolves quickly and software claims are often drafted broadly and abstractly, so conceptually identical work is described in different terms. Keyword search matches surface form and misses these. Semantic search retrieves by meaning, which is what the task requires.
How does semantic search improve AI prior art search?
Semantic search improves AI prior art search by retrieving conceptually relevant disclosures regardless of terminology, across both patents and scientific literature, and by handling the cross-domain case where relevant art is in a different field. It encodes meaning rather than surface form. Combined with an ontology, it returns a structured, high-recall view of the prior art.
Does AI prior art search need to cover scientific literature?
AI prior art search needs to cover scientific literature because the most relevant and most recent AI disclosures appear there first, in preprints, conference proceedings, and technical reports. Covering patents alone leaves the state of the art unretrieved. Cypris searches both across more than 500 million patents and scientific papers.
Which teams run AI prior art search?
AI prior art search is run by IP, R&D, and patent teams at technology companies and across industries adopting AI, as well as by patent professionals assessing novelty. It is increasingly important as AI patenting grows. Cypris serves hundreds of enterprise customers across research-intensive and regulated industries.
How current does AI prior art search need to be?
AI prior art search needs to be continuously current, because AI research and filings publish constantly and the state of the art shifts quickly. A one-time search reflects only the moment it was run. Cypris uses Agentic Monitoring to track a technology area and surface newly published disclosures as they appear.
Endnotes
- Stanford Institute for Human-Centered Artificial Intelligence (2025). Artificial Intelligence Index Report 2025, Chapter 1. arXiv:2504.07139. https://doi.org/10.48550/arxiv.2504.07139
- World Intellectual Property Organization (2024). Patent Landscape Report: Generative Artificial Intelligence. Geneva: WIPO. https://doi.org/10.34667/tind.49740
- Cavallucci, N., Chibane, I. & Ayaou, M. (2026). DAPFAM: A Domain-Aware Family-level Dataset to benchmark cross-domain patent retrieval. Array. https://doi.org/10.1016/j.array.2026.100720
- Lupu, M. (2013). Patent Retrieval. Foundations and Trends in Information Retrieval. https://doi.org/10.1561/1500000027
- Stamatis, V. (2022). End to End Neural Retrieval for Patent Prior Art Search. Lecture Notes in Computer Science. https://doi.org/10.1007/978-3-030-99739-7_66
- Zihayat, M. & Etwaroo, R. (2021). A non-factoid question answering system for prior art search. Expert Systems with Applications. https://doi.org/10.1016/j.eswa.2021.114910

Tightening regulation of per- and polyfluoroalkyl substances is reshaping materials chemistry, and it is opening patent white space for organizations that can develop fluorine-free alternatives. PFAS are used for water, oil, and stain resistance across coatings, textiles, firefighting foams, membranes, semiconductors, and food packaging, and they are now the subject of the broadest chemical restriction ever proposed in Europe. The universal PFAS restriction proposal submitted to the European Chemicals Agency in January 2023 by five national authorities covers on the order of 10,000 substances, and it drew more than 5,600 comments from over 4,400 organizations, an unprecedented response that reflects how many industries are affected.¹ The scope depends on definition: under the 2021 OECD definition, which classifies a substance as PFAS if it contains at least one fully fluorinated carbon, several million catalogued substances qualify, while the number in active commercial use is far smaller.²
The regulatory trajectory is a sequence of tightening actions rather than a single event, which is what makes the resulting innovation demand durable. In the European Union, restrictions moved from PFOS in 2006 to PFOA and related substances in later years, to a PFHxA restriction adopted in 2024, a ban on PFAS in firefighting foams, and a ban on PFAS in food-contact packaging taking effect in 2026, with the universal restriction proposal under scientific evaluation through 2026.¹ In the United States, the Environmental Protection Agency finalized the first national drinking-water limits for several PFAS in 2024, setting maximum contaminant levels of 4.0 parts per trillion for PFOA and PFOS and higher limits for other compounds, and designated PFOA and PFOS as hazardous substances under the federal cleanup statute the same year.³ ECHA has estimated that, absent action, several million tonnes of PFAS would reach the environment over the coming decades.¹
This regulatory pressure is a well-understood driver of innovation. The Porter hypothesis, that well-designed environmental regulation can induce innovation that partly or wholly offsets compliance costs, has been supported across two decades of evidence and a multi-country meta-analysis, and firm-level studies show environmental regulation inducing greener product innovation specifically in chemical industries.⁴,⁵,⁶ For materials developers, the implication is direct: regulation is converting fluorine-free chemistry from a niche into a competitive frontier, and the organizations that build defensible IP positions early will hold advantage as substitution accelerates.
Where the white space is
Firefighting foams. Fluorine-free foams are the most advanced substitution area, driven by bans on PFAS-containing aqueous film-forming foams, though performance and toxicity gaps relative to legacy foams remain an active research and patenting frontier.⁷
Textile and coating treatments. Water- and oil-repellent finishes are a major PFAS use, and fluorine-free hydrophobic and oleophobic coatings, including bio-based and hierarchical-structured approaches, are an active area of development with room for defensible positions.⁸,⁹
Membranes and packaging. Food-contact packaging faces near-term bans, and membrane and barrier applications require substitutes that match performance, which keeps white space open where a fluorine-free chemistry can meet the functional requirement.
Semiconductors and specialty uses. Certain high-performance uses have few current substitutes, so these areas are simultaneously the hardest to displace and the most valuable to solve, and the patent landscape around viable alternatives is comparatively sparse.
How to find PFAS-alternative white space
Scope the application area and functional requirement precisely, since PFAS substitution is application-specific and a fluorine-free chemistry that works for textiles may not work for firefighting foam.
Map patents and scientific literature across the fluorine-free chemistries relevant to that application, because materials research precedes patenting and gives the earliest signal of a viable alternative.
Cluster activity by concept and attribute it to organizations, using an ontology to group related chemistry and normalize assignees, so dense and sparse regions are visible.
Identify the sparse, defensible regions, distinguishing genuine white space from areas that are sparse only because a chemistry does not yet meet the functional requirement.
Monitor continuously, tracking both the chemistry and the regulatory timeline, so filings and restrictions are surfaced as they publish and a white space position is secured before substitution accelerates.
Where Cypris fits
Cypris runs patent landscape and white space analysis for regulation-driven fields such as PFAS alternatives across a corpus of more than 500 million patents and scientific papers, organized through a proprietary R&D ontology. Semantic search across patents and scientific literature surfaces fluorine-free chemistry regardless of nomenclature and connects filings to the underlying materials research, which is where the earliest signals of viable alternatives appear. The ontology clusters activity by application and chemistry and normalizes organizations to canonical entities, so a team can resolve which fluorine-free approaches are crowded and which remain open as white space. Cypris Q, the platform's agentic layer, lets teams run landscape and white space analysis conversationally and chain the search, attribution, and gap analysis, and Agentic Monitoring tracks a defined chemistry over time and flags new filings as they publish. Cypris provides enterprise API partnerships with OpenAI, Anthropic, and Google, and is built with enterprise-grade security. Cypris serves hundreds of enterprise customers across pharmaceuticals, chemicals, advanced materials, energy, and other regulated industries.
FAQ
Why is PFAS regulation creating patent white space?
PFAS regulation is creating patent white space by driving demand for fluorine-free alternatives across many applications, faster than defensible IP positions have been established. The EU REACH universal restriction covers roughly 10,000 substances and US EPA rules now limit several PFAS, so substitution is accelerating. Organizations that build fluorine-free IP early can hold advantage as demand rises.
What is the EU REACH universal PFAS restriction?
The EU REACH universal PFAS restriction is a proposal submitted to the European Chemicals Agency in January 2023 by five national authorities to restrict the manufacture and use of PFAS as a class, covering on the order of 10,000 substances. It drew more than 5,600 comments from over 4,400 organizations. It is under scientific evaluation, with the outcome expected to shape substitution across many industries.
What US rules apply to PFAS?
In the United States, the Environmental Protection Agency finalized the first national drinking-water limits for several PFAS in 2024, setting maximum contaminant levels of 4.0 parts per trillion for PFOA and PFOS and higher limits for other compounds, and designated PFOA and PFOS as hazardous substances under the federal cleanup statute the same year. These actions increase the pressure to substitute PFAS. They apply alongside state-level restrictions.
Which application areas have the most PFAS-alternative white space?
Firefighting foams, textile and coating treatments, membranes and packaging, and certain semiconductor and specialty uses all have PFAS-alternative white space, though the amount varies. Firefighting-foam alternatives are the most advanced, while high-performance specialty uses have few substitutes and are the most valuable to solve. White space is largest where a fluorine-free chemistry can meet the functional requirement but few patents yet exist.
How does regulation drive innovation in materials?
Regulation drives innovation in materials by creating demand for compliant substitutes, a pattern described by the Porter hypothesis and supported by two decades of evidence and firm-level studies in chemical industries. Well-designed regulation induces innovation that can partly offset compliance costs. For PFAS, this is converting fluorine-free chemistry from a niche into a competitive frontier.
How do you find white space in PFAS alternatives?
Finding white space in PFAS alternatives means scoping a specific application and functional requirement, mapping patents and scientific literature across the relevant fluorine-free chemistries, clustering activity by concept, and identifying the sparse, defensible regions. Because materials research precedes patenting, literature coverage gives early signal. The analysis must distinguish genuine white space from areas that are sparse because no chemistry yet meets the requirement.
Why does PFAS-alternative analysis need scientific literature?
PFAS-alternative analysis needs scientific literature because fluorine-free chemistries appear in research before they are patented, so the literature gives the earliest signal of a viable alternative. Analyzing patents alone gives a lagging view. Cypris analyzes both across more than 500 million patents and scientific papers.
Which teams work on PFAS alternatives?
PFAS alternatives are developed by R&D, innovation, and IP teams in chemicals, advanced materials, coatings, textiles, consumer products, and their suppliers, alongside regulatory affairs. The work is driven by tightening regulation and customer demand for fluorine-free products. Cypris serves hundreds of enterprise customers across chemicals, advanced materials, and other regulated industries.
How do you keep a PFAS-alternatives landscape current? Keeping a PFAS-alternatives landscape current requires continuous monitoring of both the chemistry and the regulatory timeline, because filings and restrictions evolve constantly. A one-time landscape ages quickly as new rules and patents publish. Cypris uses Agentic Monitoring to track a defined chemistry over time and flag new filings as they publish.
Endnotes
- European Chemicals Agency. Registry of restriction intentions: per- and polyfluoroalkyl substances (PFAS) universal restriction proposal (2023) and related consultation and evaluation materials. https://echa.europa.eu/
- OECD (2021). Reconciling Terminology of the Universe of Per- and Polyfluoroalkyl Substances: Recommendations and Practical Guidance; and US Environmental Protection Agency PFAS inventory materials.
- US Environmental Protection Agency (2024). PFAS National Primary Drinking Water Regulation; and CERCLA designation of PFOA and PFOS as hazardous substances. https://www.epa.gov/pfas
- Ambec, S., Cohen, M. A., Elgie, S. & Lanoie, P. (2013). The Porter Hypothesis at 20. Review of Environmental Economics and Policy. https://doi.org/10.1093/reep/res016
- Yan, Z., Li, Y., Zhang, X. & Zhu, J. (2024). Revisiting the Porter hypothesis: a multi-country meta-analysis. Humanities and Social Sciences Communications. https://doi.org/10.1057/s41599-024-02671-9
- Choi, J., Kang, J. & Chung, S. (2025). Environmental regulation, induced innovation, and greener transition: firm-level evidence. Journal of Development Economics. https://doi.org/10.1016/j.jdeveco.2025.103678
- Hossain, T., Ormond, R. B. et al. (2024). Exploring the Prospects and Challenges of Fluorine-Free Firefighting Foams (F3) as Alternatives to AFFF: A Review. ACS Omega. https://doi.org/10.1021/acsomega.4c03673
- Likozar, B. et al. (2024). Unveiling PFAS-free Solutions for Hydrophobic and Oleophobic Textile Coatings. https://doi.org/10.55295/psl.2024.i19
- Nicolas, M. et al. (2024). PFAS-free hierarchical superhydrophobic textiles. Advanced Engineering Materials. https://doi.org/10.1002/adem.202401736

Freedom-to-operate for GLP-1 receptor agonists and peptide therapeutics is among the most demanding FTO problems in pharmaceuticals, because protection in this class is built as a dense, layered thicket that extends far beyond the active ingredient. Freedom-to-operate determines whether making, using, or selling a product would infringe another party's active patent claims. In the GLP-1 and peptide space, answering that question requires reading many claim types across many patents, because a single product is protected by a stack of filings covering the molecule, its formulation, its dosing, its delivery device, and its manufacture. A peer-reviewed analysis of GLP-1 receptor agonists approved between 2005 and 2021 found that manufacturers listed a median of 19.5 patents per product, that 54 percent of those patents were on delivery devices rather than the active ingredient, that the median expected protection was 18.3 years after approval, and that no generic manufacturer had yet successfully challenged a GLP-1 receptor agonist patent.¹
The commercial stakes are large. Industry analyst forecasts vary widely with scope, placing the GLP-1 market anywhere from the low tens of billions of dollars to well over one hundred billion by 2030 and projecting double-digit annual growth; these are analyst estimates rather than authoritative figures, and they differ mainly in what they count.² The scale of the opportunity is what drives the density of the patent thicket, because each additional protected feature can delay competition on a high-revenue product. For any organization developing a follow-on peptide, a biosimilar, or a differentiated GLP-1 product, FTO is therefore a gating analysis rather than a formality.
Peptide therapeutics compound the difficulty. Peptides can be claimed as sequences and modifications, formulated for stability and half-life extension, delivered by injection or increasingly by oral routes, and manufactured through distinct synthesis and purification processes, so the claim surface is broad. Recent filing activity has shifted toward oral delivery, dual and triple receptor agonists, and combination therapies, which is where both the newest FTO risk and the remaining white space now sit.³ An FTO analysis in this class has to cover all of these dimensions, and it has to stay current as the frontier moves.
What creates FTO risk in GLP-1 and peptide products
Composition-of-matter claims. These cover the peptide itself, including sequences, analogues, and modifications, and are the primary protection, though in a mature class many core molecules approach expiry.
Formulation claims. These cover stabilized, extended-release, and oral formulations, which are heavily patented, as formulation is where much peptide innovation and differentiation occurs.
Dosing-regimen and method-of-use claims. These cover titration schedules and specific therapeutic uses, and can block a product for a particular indication or regimen even when the molecule is otherwise available.
Delivery-device claims. These cover injection pens and other devices and are a large share of the thicket; peer-reviewed analysis found delivery devices accounted for the majority of listed GLP-1 patents and function as a distinct barrier to entry.¹,⁴
Process and manufacturing claims. These cover synthesis and purification routes, so a developer can be free to use a molecule yet blocked from a particular manufacturing method.
A single molecule illustrates the layering. A published patent landscape of one dual GLP-1/glucagon receptor agonist identified twelve patent families spanning composition-of-matter, process chemistry, formulation, dosing regimen, and method-of-use, a clean worked example of how all five claim types stack on one product.⁵
The thicket dynamic and the expiry landscape
The density of GLP-1 protection reflects a broader pharmaceutical pattern. Empirical analysis shows the number of patents filed per active ingredient rose from 1.86 in 2001 to nearly six by 2019, driven substantially by continuation applications, which account for roughly a third of small-molecule pharmaceutical patents.⁶ These secondary filings extend the effective protection period, and the economics of that extension, including how patent challenges and settlements shape effective market life, are well documented.⁷ Pharmaceutical thickets also differ structurally from thickets in complex-technology industries, which is why FTO methods developed for electronics do not transfer cleanly to peptides.⁸
The expiry landscape is the other half of the picture. As core molecules approach the end of composition-of-matter protection, the surrounding formulation, device, and process claims determine when and where competition can actually enter. Analysis of one leading GLP-1 molecule found that the timing of primary-patent expiry varies substantially by market, so freedom-to-operate for a follow-on product is jurisdiction-specific, and the practical entry date is governed by the secondary thicket rather than the headline molecule expiry.⁹ For a developer, this means FTO must be assessed claim-by-claim and market-by-market, not at the level of the molecule.
How AI-powered FTO helps
Navigating a thicket of this density by manual search is slow and prone to coverage gaps, which are the main source of FTO risk. AI-powered FTO addresses this with semantic search that retrieves relevant claims regardless of terminology, claim-level analysis that focuses on the independent claims defining infringement scope across all five claim types, and continuous monitoring that keeps a cleared position current as new formulation, device, and combination filings publish. Because peptide innovation appears in scientific literature before it is patented, reading both patents and literature gives earlier warning of where the thicket is extending.
Where Cypris fits
Cypris runs claim-level, semantic, AI-powered freedom-to-operate across a corpus of more than 500 million patents and scientific papers, organized through a proprietary R&D ontology. Semantic search across patents and scientific literature surfaces relevant claims regardless of terminology, across composition, formulation, dosing-regimen, delivery-device, and process claims, which is what a dense peptide thicket demands. The ontology clusters the thicket by concept and normalizes assignees, so a team sees the structure of protection around a molecule rather than a flat list. Cypris Q, the platform's agentic layer, lets teams run and chain FTO analysis conversationally, and Agentic Monitoring tracks a molecule and its surrounding thicket over time, flagging new formulation, device, and combination filings as they publish. Cypris provides enterprise API partnerships with OpenAI, Anthropic, and Google, and is built with enterprise-grade security. Cypris serves hundreds of enterprise customers across pharmaceuticals, chemicals, advanced materials, energy, and other regulated industries.
FAQ
Why is freedom-to-operate hard for GLP-1 and peptide therapeutics?
Freedom-to-operate is hard for GLP-1 and peptide therapeutics because protection is built as a dense, layered thicket extending well beyond the active ingredient. A peer-reviewed analysis found GLP-1 products carry a median of 19.5 listed patents each, most of them on delivery devices. Assessing FTO requires reading composition, formulation, dosing, device, and process claims across many patents and markets.
What claim types create FTO risk for GLP-1 products?
Five claim types create FTO risk for GLP-1 products: composition-of-matter claims on the peptide, formulation claims on stabilized and oral forms, dosing-regimen and method-of-use claims, delivery-device claims, and process or manufacturing claims. Each can independently block a product. Delivery-device claims are a particularly large share of the GLP-1 thicket.
How many patents protect a typical GLP-1 product?
A peer-reviewed analysis of GLP-1 receptor agonists approved between 2005 and 2021 found a median of 19.5 listed patents per product, with 54 percent on delivery devices rather than the active ingredient, and a median of 18.3 years of expected protection after approval. No generic manufacturer had successfully challenged a GLP-1 receptor agonist patent as of that analysis. These figures illustrate the density of the thicket.
What is a pharmaceutical patent thicket?
A pharmaceutical patent thicket is a dense set of overlapping patents around a single product that extends protection beyond the core molecule. Empirical analysis shows patents per active ingredient rose from 1.86 in 2001 to nearly six by 2019, driven substantially by continuation applications. Thickets shape when and where competition can enter.
How does the expiry of GLP-1 patents affect freedom-to-operate?
The expiry of GLP-1 patents affects freedom-to-operate market-by-market, because primary-patent expiry timing varies by jurisdiction and the practical entry date is governed by the surrounding formulation, device, and process claims rather than the molecule alone. FTO must therefore be assessed claim-by-claim and market-by-market. A molecule can be off-patent in one country and still protected in another.
Where is the white space in GLP-1 and peptide development?
Recent filing activity has shifted toward oral delivery, dual and triple receptor agonists, and combination therapies, which is where both new FTO risk and remaining white space now sit. Mapping this frontier requires reading patents and scientific literature together, since peptide innovation appears in research first. White space analysis identifies the areas that are still open.
How does AI-powered FTO help with peptide therapeutics?
AI-powered FTO helps with peptide therapeutics by using semantic search to retrieve relevant claims regardless of terminology, claim-level analysis to focus on the independent claims that define infringement across all claim types, and continuous monitoring to keep a cleared position current. This is what a dense, fast-moving thicket requires. Cypris runs this across more than 500 million patents and scientific papers.
Which teams need GLP-1 and peptide FTO analysis?
GLP-1 and peptide FTO analysis is needed by pharmaceutical and biotech R&D, IP, and business-development teams developing follow-on peptides, biosimilars, differentiated formulations, or combination products. It is also relevant to generics manufacturers assessing entry. Cypris serves hundreds of enterprise customers across pharmaceuticals and other regulated industries.
How current does GLP-1 FTO need to be?
GLP-1 FTO needs to be continuously current, because new formulation, device, dosing, and combination filings publish constantly and can change a cleared position. A one-time assessment reflects only the moment it was run. Cypris uses Agentic Monitoring to track a molecule and its surrounding thicket over time and flag new filings as they publish.
Endnotes
- Tu, S. S., Feldman, W. B., Alhiary, R., Gabriele, S., Kesselheim, A. S. & Beall, R. F. (2023). Patents and Regulatory Exclusivities on GLP-1 Receptor Agonists. JAMA. https://doi.org/10.1001/jama.2023.13872
- Industry analyst estimates (e.g., Research and Markets; BCC Research). GLP-1 market forecasts vary widely by scope and are presented here as order-of-magnitude estimates, not authoritative figures.
- Han, J., Zhou, Z., Jiang, N. & Lu, W. (2023). An updated patent review of GLP-1 receptor agonists (2020–present). Expert Opinion on Therapeutic Patents. https://doi.org/10.1080/13543776.2023.2274905
- Tu, S. S., Feldman, W. B. et al. (2024). Delivery Device Patents on GLP-1 Receptor Agonists. JAMA. https://doi.org/10.1001/jama.2024.0919
- Fasi, M. A. (2026). Patent landscape and therapeutic evolution of mazdutide. Expert Opinion on Therapeutic Patents. https://doi.org/10.1080/13543776.2026.2645812
- Tu, S. S. (2024). The Long CON: An Empirical Analysis of Pharmaceutical Patent Thickets. University of Pittsburgh Law Review. https://doi.org/10.5195/lawreview.2024.1049
- Hemphill, C. S. & Sampat, B. N. (2012). Evergreening, patent challenges, and effective market life in pharmaceuticals. Journal of Health Economics. https://doi.org/10.1016/j.jhealeco.2012.01.004
- Tu, S. S. & Carrier, M. A. (2023). Why Pharmaceutical Patent Thickets Are Unique. SSRN. https://doi.org/10.2139/ssrn.4571486
- Ramesh, S., Cross, S., Levi, J., Hill, A. & Venter, F. (2026). How Low Could Semaglutide Prices Fall? Implications for Global Access Ahead of Patent Expiry. Obesity. https://doi.org/10.1002/oby.70241
.avif)
