A faster, more accurate way to explore innovation data—now available in Cypris.
For innovation teams, speed and accuracy aren’t optional—they’re critical. You need to quickly find all relevant documents, slice and dice datasets however you want, and trust that the results are complete and representative. With this in mind, we’ve upgraded how semantic search works inside Cypris.
Today, we’re launching an upgraded search infrastructure that gives users access to full, exact result sets—unlocking more powerful analysis, faster iteration, and deterministic filtering and charting.
Unlike traditional semantic or vector search engines—which make it difficult to count, filter, or chart large sets of matched documents—our new approach prioritizes transparency and performance while preserving semantic relevance.
Why we moved away from vector search
Our original implementation relied on semantic and vector search to capture the “meaning” behind user queries. But as our platform evolved, it became clear that these systems weren’t well-suited for our core use cases.
Users needed:
- Deterministic filtering (e.g., "how many results match this atom?")
- Transparent, complete result sets to power charts and dashboards
- Fast, repeatable queries that don’t change subtly over time
Modern vector search systems don’t easily support this level of transparency. They return approximate matches and abstract similarity scores, often making it hard to understand why a document was returned—or whether it’s the full picture.
So we made a decision: move away from vector search and lean into what traditional search engines do best.
A return to boolean and lexical search—with a twist
We rebuilt our search infrastructure on top of Elasticsearch’s powerful boolean and lexical search capabilities. This shift brings major advantages:
- Faster query speeds that dramatically improve iteration time
- Deterministic filtering and counts, so every chart is grounded in the full dataset
- Predictable, explainable results that users can trust
But we didn’t stop there.
To preserve the benefits of semantic understanding, we’ve rethought where that intelligence should live—not at query time, but at data ingestion.
Capturing semantic meaning at ingest time
Instead of computing document-query similarity during search, we enrich documents at the time of ingestion. Here’s how:
- Synonym expansion: We find related words and concepts not explicitly mentioned in the document and add them as fields, enabling semantic-style recall via lexical search.
- Stemming: Both queries and documents are reduced to their root forms, allowing consistent matches (e.g., “running” and “run”).
The result? You get the same functionality—semantically relevant results—without the opacity or latency tradeoffs of vector search.
What’s next: Reranking for even better relevance
We’re not done. Coming soon to Cypris is a reranking layer that boosts the most relevant results to the top of the list using lightweight vector techniques.
Here’s how it works:
- A standard lexical search retrieves the full result set.
- We take the top N results and rerank them using vector similarity, powered by Elasticsearch’s new hybrid scoring capabilities.
- You get faster queries with even better relevance—without compromising on counts or transparency.
This layered approach gives us the best of both worlds: precise filtering and fast queries, plus smarter ordering of results where it matters most.
We’re excited to bring this upgrade to our users, and we’re already seeing teams iterate faster and uncover insights more confidently. This is a foundational shift—and just the beginning of what’s to come.
Want a walkthrough of what’s changed? Reach out to our team.

Introducing our upgraded semantic search
A faster, more accurate way to explore innovation data—now available in Cypris.
For innovation teams, speed and accuracy aren’t optional—they’re critical. You need to quickly find all relevant documents, slice and dice datasets however you want, and trust that the results are complete and representative. With this in mind, we’ve upgraded how semantic search works inside Cypris.
Today, we’re launching an upgraded search infrastructure that gives users access to full, exact result sets—unlocking more powerful analysis, faster iteration, and deterministic filtering and charting.
Unlike traditional semantic or vector search engines—which make it difficult to count, filter, or chart large sets of matched documents—our new approach prioritizes transparency and performance while preserving semantic relevance.
Why we moved away from vector search
Our original implementation relied on semantic and vector search to capture the “meaning” behind user queries. But as our platform evolved, it became clear that these systems weren’t well-suited for our core use cases.
Users needed:
- Deterministic filtering (e.g., "how many results match this atom?")
- Transparent, complete result sets to power charts and dashboards
- Fast, repeatable queries that don’t change subtly over time
Modern vector search systems don’t easily support this level of transparency. They return approximate matches and abstract similarity scores, often making it hard to understand why a document was returned—or whether it’s the full picture.
So we made a decision: move away from vector search and lean into what traditional search engines do best.
A return to boolean and lexical search—with a twist
We rebuilt our search infrastructure on top of Elasticsearch’s powerful boolean and lexical search capabilities. This shift brings major advantages:
- Faster query speeds that dramatically improve iteration time
- Deterministic filtering and counts, so every chart is grounded in the full dataset
- Predictable, explainable results that users can trust
But we didn’t stop there.
To preserve the benefits of semantic understanding, we’ve rethought where that intelligence should live—not at query time, but at data ingestion.
Capturing semantic meaning at ingest time
Instead of computing document-query similarity during search, we enrich documents at the time of ingestion. Here’s how:
- Synonym expansion: We find related words and concepts not explicitly mentioned in the document and add them as fields, enabling semantic-style recall via lexical search.
- Stemming: Both queries and documents are reduced to their root forms, allowing consistent matches (e.g., “running” and “run”).
The result? You get the same functionality—semantically relevant results—without the opacity or latency tradeoffs of vector search.
What’s next: Reranking for even better relevance
We’re not done. Coming soon to Cypris is a reranking layer that boosts the most relevant results to the top of the list using lightweight vector techniques.
Here’s how it works:
- A standard lexical search retrieves the full result set.
- We take the top N results and rerank them using vector similarity, powered by Elasticsearch’s new hybrid scoring capabilities.
- You get faster queries with even better relevance—without compromising on counts or transparency.
This layered approach gives us the best of both worlds: precise filtering and fast queries, plus smarter ordering of results where it matters most.
We’re excited to bring this upgrade to our users, and we’re already seeing teams iterate faster and uncover insights more confidently. This is a foundational shift—and just the beginning of what’s to come.
Want a walkthrough of what’s changed? Reach out to our team.

Keep Reading

Electrolysis has become the center of gravity in hydrogen innovation, and the electrolyzer patent landscape is where the clean-hydrogen transition is being contested. A joint study of global patent data by the European Patent Office and the International Energy Agency found that technologies motivated by climate concerns accounted for nearly 80 percent of all hydrogen-production patents by 2020, with growth driven chiefly by a sharp increase in innovation in water electrolysis, and that climate-driven hydrogen technologies generated roughly twice as many international patent families as established, fossil-based methods.¹ The commercial backdrop is a projected expansion of electrolyzer manufacturing on the order of a 65-fold increase in market size over the decade, as countries scale low-emissions hydrogen for hard-to-abate sectors.²,³ For R&D and IP teams, the strategic questions are which electrolyzer technology route to back and where defensible IP positions remain, and both are patent-landscape questions.
The landscape divides across four electrolyzer technologies at different maturity levels, each a distinct region of patenting, and each characterized in the US Department of Energy's comparative assessment of solid-oxide, alkaline, and proton-exchange-membrane electrolyzers.⁴ Alkaline electrolysis is the most mature and lowest-cost route, using a liquid alkaline electrolyte and avoiding scarce precious metals, so its patenting concentrates on efficiency, dynamic operation to follow variable renewable power, and stack scale-up. Proton-exchange-membrane (PEM) electrolysis offers compact, responsive operation well suited to variable renewables but relies on scarce platinum-group catalysts and specialized membranes, so a large share of its patenting targets catalyst loading reduction, membrane durability, and cost.⁵ Solid-oxide electrolysis (SOEC) operates at high temperature with high electrical efficiency and can co-electrolyze to produce syngas, but durability and thermal cycling are the central challenges, so patenting concentrates there. Anion-exchange-membrane (AEM) electrolysis is the newest route, aiming to combine PEM-like performance without precious-metal dependence, and it is the least mature and least crowded, which makes it a notable area of white space; its membranes and non-precious-metal catalysts are an active peer-reviewed research frontier.⁶
Geography and institutional origin further shape the landscape. The EPO and IEA analysis found Europe gaining an edge as a location for electrolyzer innovation and manufacturing investment, while Japan led patenting in hydrogen end-use for the automotive sector, and it noted that momentum in other end-use applications, such as aviation, shipping, and power generation, had not yet matched the attention those sectors receive.¹ The European Commission's Joint Research Centre has separately tracked the status of water electrolysis and hydrogen technology in the European Union, corroborating the region's manufacturing push.⁷ It also found that emerging low-emissions hydrogen carriers, including liquid organic hydrogen carriers and ammonia cracking, grew (by about 12.5 percent and 7.8 percent in international patent families respectively) with roughly half of that activity originating in universities and public research, an early-stage signal of where future commercial IP may form.¹ Because applications publish about eighteen months after filing, the most recent activity, particularly in the newer AEM and SOEC routes, is under-represented, so the current frontier is more active than granted-patent counts suggest.
The four electrolyzer routes and where white space sits
Alkaline. The most mature and lowest-cost route, avoiding precious metals; patenting concentrates on efficiency, dynamic operation, and scale-up, so it is comparatively crowded on core design.
PEM. Compact and responsive but reliant on platinum-group catalysts and specialized membranes; white space centers on catalyst reduction, membrane durability, and cost.
SOEC. High-temperature and high-efficiency with co-electrolysis potential, but durability and thermal cycling are the open problems where patenting and white space concentrate.
AEM. The newest route, aiming for PEM-like performance without precious metals; the least mature and least crowded, and therefore a notable area of white space.²
Carriers and end-use. Liquid organic hydrogen carriers and ammonia cracking are early-stage and university-driven, and several end-use sectors beyond automotive remain comparatively under-patented.¹
How AI-powered landscape and white space analysis helps
Resolving four technology routes at different maturities, across geographies and institutions, requires more than keyword search. AI-powered analysis addresses this with semantic search that clusters activity by route and by the problem being solved across varied terminology, attribution that normalizes filers to canonical entities and distinguishes university from commercial activity, and continuous monitoring that tracks the newer routes where recent activity is under-represented. Because electrolyzer advances appear in scientific literature before they are patented, reading both patents and literature gives the earliest signal of where the frontier and the white space are moving.
Where Cypris fits
Cypris runs patent landscape and white space analysis for multi-route energy fields such as hydrogen electrolysis across a corpus of more than 500 million patents and scientific papers, organized through a proprietary R&D ontology. The ontology clusters activity by electrolyzer route, alkaline, PEM, SOEC, and AEM, and by the problem being solved, and normalizes filers to canonical entities, so a team can resolve which routes and problems are crowded and which, such as AEM and SOEC durability, remain open as white space. Semantic search across patents and scientific literature connects filings to the underlying materials and engineering research, which is where electrolyzer advances appear first, and distinguishes university from commercial activity. Cypris Q, the platform's agentic layer, lets teams run landscape and white space analysis conversationally and chain the clustering, attribution, and gap analysis, and Agentic Monitoring tracks a defined route over time and flags new patents and papers as they publish, which is essential where the newest routes are under-represented by publication lag. Cypris provides enterprise API partnerships with OpenAI, Anthropic, and Google, and is built with enterprise-grade security. Cypris serves hundreds of enterprise customers across pharmaceuticals, chemicals, advanced materials, energy, and other regulated industries.
FAQ
Why is electrolysis the focus of hydrogen patenting?
Electrolysis is the focus of hydrogen patenting because it can produce hydrogen with zero direct emissions when powered by renewable or nuclear electricity. A joint EPO and IEA study found that climate-motivated technologies accounted for nearly 80 percent of hydrogen-production patents by 2020, with growth driven chiefly by a surge in electrolysis. Climate-driven hydrogen technologies generated roughly twice the international patent families of established methods.
What are the main electrolyzer technologies?
The main electrolyzer technologies are alkaline, proton-exchange-membrane (PEM), solid-oxide (SOEC), and anion-exchange-membrane (AEM). They differ in maturity, cost, materials, and operating conditions, and each occupies a distinct region of the patent landscape. Alkaline is the most mature and AEM the newest.
Where is the white space in the electrolyzer patent landscape?
The white space in the electrolyzer patent landscape is concentrated in anion-exchange-membrane electrolysis, which is the newest and least crowded route, in solid-oxide durability and thermal cycling, in reducing precious-metal catalyst use and improving membrane durability in PEM, and in early-stage hydrogen carriers such as liquid organic carriers and ammonia cracking. Core alkaline design is comparatively crowded. The higher-value opportunities are in the newer routes and unsolved durability problems.
How do the electrolyzer routes trade off?
The electrolyzer routes trade off maturity, cost, and materials. Alkaline is mature and low-cost but less dynamic; PEM is responsive but relies on scarce platinum-group metals; SOEC is highly efficient but faces durability challenges; and AEM aims to combine PEM-like performance without precious metals but is the least mature. Each route's patenting concentrates on its specific weakness.
How fast is the electrolyzer market expected to grow?
The electrolyzer market is expected to grow rapidly, with the IEA projecting an expansion on the order of a 65-fold increase in market size over the decade as countries scale low-emissions hydrogen. This growth is the commercial driver behind the surge in electrolysis patenting. It also raises the value of securing defensible IP positions early.
Which regions lead electrolyzer innovation?
The EPO and IEA analysis found Europe gaining an edge as a location for electrolyzer innovation and manufacturing investment, while Japan led hydrogen end-use patenting in the automotive sector. Momentum in several other end-use sectors had not yet matched the attention they receive. The geographic distribution differs by technology route and end-use.
Why does electrolyzer analysis need scientific literature?
Electrolyzer analysis needs scientific literature because materials and engineering advances, particularly in catalysts, membranes, and the newer routes, appear in research before they are patented, so the literature gives the earliest signal. Analyzing patents alone gives a lagging view, and much early activity is university-driven. Cypris analyzes both across more than 500 million patents and scientific papers.
Which teams use electrolyzer patent landscape analysis?
Electrolyzer patent landscape analysis is used by R&D, innovation, IP, and strategy teams at electrolyzer and equipment makers, energy and industrial-gas companies, materials developers, and their partners, as well as investors. It informs which route to back, where to file, and where competitors are concentrated. Cypris serves hundreds of enterprise customers across energy, advanced materials, chemicals, and other regulated industries.
How do you keep an electrolyzer landscape current?
Keeping an electrolyzer landscape current requires continuous monitoring, because the field moves quickly, the newer routes are advancing, and publication lag under-represents the most recent activity. A one-time landscape ages quickly. Cypris uses Agentic Monitoring to track a defined route and flag new patents and papers as they publish.
Endnotes
- European Patent Office & International Energy Agency (2023). Hydrogen patents for a clean energy future: A global trend analysis of innovation along hydrogen value chains. https://www.iea.org/reports/hydrogen-patents-for-a-clean-energy-future
- International Energy Agency, reported via World Economic Forum (2023). Hydrogen patent filings: Europe and Japan lead on innovation (projected ~65-fold electrolyzer market growth this decade). https://www.weforum.org/stories/2023/03/hydrogen-innovation-patents-technology/
- International Energy Agency. Global Hydrogen Review (annual series). https://www.iea.org/reports/global-hydrogen-review-2024
- Kelly, J. C., Elgowainy, A. & Iyer, R. (2022). Electrolyzers for Hydrogen Production: Solid Oxide, Alkaline, and Proton Exchange Membrane. US Department of Energy (OSTI). https://www.osti.gov/
- US Department of Energy (2024). Hydrogen Shot: Water Electrolysis Technology Assessment. https://www.energy.gov/
- Zhang, M. et al. (2024). Advanced development of anion-exchange membrane electrolyzers for hydrogen production: from anion-exchange membranes to membrane electrode assemblies. Chemical Communications. https://doi.org/10.1039/D3CC05904A
- European Commission Joint Research Centre (2023). Water electrolysis and hydrogen in the European Union: Status Report on Technology Development, Trends, Value Chains and Markets. https://publications.jrc.ec.europa.eu/

Energy storage is one of the fastest-growing domains of patenting. A joint analysis by the International Energy Agency and the European Patent Office found that patenting in batteries and electricity storage grew at an average of 14 percent per year between 2005 and 2018, roughly four times faster than the all-technology average, across more than 65,000 international patent families, with batteries accounting for the large majority of electricity-storage patenting.¹ More recent IEA analysis reports that batteries have come to dominate the energy patent landscape.² The drivers are structural: the electrification of transport, the decarbonization of the grid, and the need for long-duration storage to balance intermittent renewable generation. These forces have pushed research and filing activity up sharply across several distinct storage technologies at once, and much of the technology that will define the market at the end of the decade is entering the patent record now.
The energy-storage landscape is not a single field but a set of competing technology routes at different technology-readiness levels, and a rigorous landscape has to segment them. Lithium-ion remains the incumbent, with filing activity concentrated on energy density, fast charging, safety, and cell-to-pack manufacturing. Solid-state batteries have seen filing activity grow several-fold since the late 2010s, and the locus of innovation has shifted from electrolyte materials discovery toward interfacial engineering and scalable manufacturing, a transition documented across recent reviews of all-solid-state commercialization.³,⁴ Within that route, the principal electrolyte classes, sulfide, oxide, polymer, and composite, present different trade-offs: sulfide solid electrolytes reach room-temperature ionic conductivities on the order of 10 to the minus three siemens per centimeter, comparable to conventional liquid electrolytes, but the dominant technical barriers are interfacial resistance, electrochemical stability at the electrode interfaces, dendrite suppression, and scalable synthesis of the electrolyte.³,⁴,⁵ Hydrogen storage, particularly solid-state routes using metal hydrides, has surged as fuel-cell and stationary applications advance, with claim activity concentrated on intermetallic alloy families and multi-phase crystal-structure engineering to balance gravimetric capacity against kinetics and operating pressure. Long-duration and grid-scale storage is an active emerging area, where vanadium redox and other flow batteries, compressed-air storage, iron-air chemistries, and thermal and gravity approaches compete, and a large share of the relevant patents are still pending.
That segmentation is the value of patent landscape and white space analysis for the energy transition. A landscape maps where filing activity concentrates, which routes and sub-classes are crowded, and which organizations are most active; a white space analysis maps where activity is sparse, revealing directions where a defensible position is still available. In a field advancing this quickly, where the architectures that will define the 2030 market are being filed today, the ability to resolve both the dense and the sparse regions, at the level of specific technology routes and sub-classes, and to track how they shift, is what converts patent data into strategic positioning.
Why energy patenting is surging
Transport electrification. The transition to electric vehicles drives intense filing in battery chemistries, energy density, fast charging, safety, and manufacturing.
Grid decarbonization. Balancing intermittent renewables requires storage, which drives filing in grid-scale and long-duration technologies.
Long-duration storage demand. Storing energy over many hours or seasonally has pushed activity into flow, compressed-air, iron-air, thermal, and hydrogen routes at differing readiness levels.
Materials and interface innovation. Much of the activity is in materials and interfaces, from solid electrolytes and metal hydrides to electrode-electrolyte engineering, where the underlying research is published before it is patented.
Publication lag. The most recent filings are under-represented because applications publish about eighteen months after their priority date, so current activity is larger than the latest figures show.
How to run an energy patent landscape and white space analysis
Scope the technology space with classification codes, selecting the relevant Cooperative Patent Classification and International Patent Classification categories for the storage routes and sub-classes in view, so the boundary is standardized and reproducible.
Aggregate to the patent-family level, so international coverage of a single invention is not double-counted and volume reflects distinct R&D.
Segment by technology route, separating lithium-ion, solid-state and its electrolyte classes, metal-hydride hydrogen storage, and the long-duration routes, since each is at a different readiness level and must be assessed on its own terms.
Cluster activity by concept using semantic analysis over classification and text, so related work groups together across the varied terminology of materials, chemistries, and architectures.
Map the dense and sparse regions and attribute activity to canonical organizations, identifying crowded sub-classes and open white space and resolving assignee variants to single entities.
Correct for publication lag and monitor continuously, discounting the most recent windows and tracking the landscape over time, because a static snapshot ages quickly in a fast-moving field.
Where Cypris fits
Cypris runs patent landscape and white space analysis for fast-moving fields such as the energy transition across a corpus of more than 500 million patents and scientific papers, organized through a proprietary R&D ontology. That structure lets Cypris segment energy-storage activity by technology route and cluster it by concept across the varied terminology of materials, chemistries, and architectures, so a team can resolve which routes and sub-classes are crowded and which remain open as white space. Dense semantic search across patents and scientific literature connects filings to the underlying materials and interface research, which matters in energy storage because the earliest signals appear in the literature before patents. Cypris Q, the platform's agentic layer, lets teams run landscape and white space analysis conversationally and chain the classification, clustering, attribution, and gap analysis. Agentic Monitoring tracks a defined storage route over time and flags new patents and papers as they publish, which is essential where recent activity is under-represented by publication lag. Cypris provides enterprise API partnerships with OpenAI, Anthropic, and Google, and is built with enterprise-grade security. Cypris serves hundreds of enterprise customers across pharmaceuticals, chemicals, advanced materials, energy, and other regulated industries.
FAQ
Why is energy storage one of the fastest-growing patent areas?
Energy storage is one of the fastest-growing patent areas because of transport electrification, grid decarbonization, and the need for long-duration storage. A joint IEA and EPO analysis found battery and electricity-storage patenting grew about 14 percent per year from 2005 to 2018, roughly four times the all-technology average, across more than 65,000 international patent families. More recent IEA analysis reports that batteries now dominate the energy patent landscape.
What technology routes does the energy-storage patent landscape cover?
The energy-storage patent landscape covers several competing routes at different readiness levels, including lithium-ion, solid-state batteries with sulfide, oxide, polymer, and composite electrolytes, metal-hydride hydrogen storage, and long-duration routes such as flow, compressed-air, iron-air, thermal, and gravity storage. Each is a distinct route with its own activity level and technical barriers. A landscape analysis segments these rather than treating storage as one field.
What is a patent landscape analysis for the energy transition?
A patent landscape analysis for the energy transition maps where filing activity concentrates across energy-storage routes, which sub-classes are crowded, and which organizations are most active, scoped by classification codes and aggregated to the patent-family level. It gives R&D and IP teams a structured, reproducible view of a fast-moving field. Paired with white space analysis, it also identifies the sparse regions where a defensible position is still available.
How do you find white space in energy-storage patents?
Finding white space in energy-storage patents means mapping patents and scientific literature across the routes, clustering activity by concept, and identifying the sparse sub-classes where few patents exist. Because materials and interface research is published before it is patented, literature coverage reveals white space earlier. The sparse regions indicate directions where a team can still build a novel, defensible position.
Why use classification codes and patent families in an energy landscape?
Classification codes scope the technology space in a standardized, reproducible way independent of applicant terminology, and patent-family aggregation avoids double-counting the multiple international applications a single invention generates. Together they make the landscape accurate and comparable across competitors. Keyword-only scoping and document-level counting distort both boundary and volume.
What are the main technical barriers in solid-state batteries?
The main technical barriers in solid-state batteries are interfacial resistance and stability at the electrode-electrolyte interfaces, dendrite suppression, and scalable synthesis and manufacturing of the solid electrolyte. Sulfide electrolytes reach ionic conductivities comparable to liquid electrolytes, so the current focus has shifted from materials discovery toward interface engineering and manufacturing. Patent activity reflects this shift.
Why does publication lag matter in energy patent landscapes? Publication lag matters because applications publish about eighteen months after their priority date, so the most recent filing activity is under-represented in current data. In a fast-moving field like energy storage, apparent softness in the latest window is usually an artifact of lag rather than a real slowdown. Longer-window trends and continuous monitoring are more reliable than the latest figures alone.
Why does energy patent analysis need scientific literature?
Energy patent analysis needs scientific literature because much of the innovation is in materials and interfaces, which are typically published in research before they are patented. Analyzing patents alone gives a lagging view, while adding literature reveals emerging activity earlier. Cypris analyzes both across more than 500 million patents and scientific papers.
How do you keep an energy patent landscape current?
Keeping an energy patent landscape current requires continuous monitoring, because the field moves quickly, new filings and research publish constantly, and publication lag hides the most recent activity. A one-time landscape ages fast. Cypris uses Agentic Monitoring to track a defined storage route over time and flag new patents and papers as they publish.
Who uses patent landscape analysis for the energy transition?
Patent landscape analysis for the energy transition is used by R&D, innovation, IP, and strategy teams at battery makers, automotive and energy companies, materials developers, and their partners. It informs where to invest, where to file, and where competitors are concentrating. Cypris serves hundreds of enterprise customers across energy, advanced materials, chemicals, and other regulated industries.
Works Cited
- International Energy Agency & European Patent Office (2020). Innovation in Batteries and Electricity Storage: A Global Analysis Based on Patent Data. https://www.iea.org/reports/innovation-in-batteries-and-electricity-storage
- International Energy Agency (2026). The State of Energy Innovation 2026. https://www.iea.org/reports/the-state-of-energy-innovation-2026
- Kim, J.-J. et al. (2026). Key Challenges and Strategies for Commercialization of All-Solid-State Batteries: Materials, Interface Engineering, and Manufacturing Processes. International Journal of Energy Research. https://doi.org/10.1155/er/8704807
- Liu, Q. et al. (2023). Interfacial Modification, Electrode/Solid-Electrolyte Engineering, and Monolithic Construction of Solid-State Batteries. Electrochemical Energy Reviews. https://doi.org/10.1007/s41918-022-00167-1
- Gamo, H., Nagai, A. & Matsuda, A. (2023). Toward Scalable Liquid-Phase Synthesis of Sulfide Solid Electrolytes for All-Solid-State Batteries. Batteries. https://doi.org/10.3390/batteries9070355
.jpg)
The Model Context Protocol has become the connective tissue between AI assistants and the specialized data that R&D and IP teams depend on. Instead of copying patent claims into a chat window or pasting abstracts from a database, a team can connect an AI client directly to patent and scientific literature sources and work in natural language. But 2026 has surfaced a sharper distinction than "which server connects to which database." The more important question for innovation leaders is whether a server is a single-source connector or a domain-oriented intelligence layer built to support the actual decisions in an R&D and IP stage-gate process. This ranked guide covers the most capable options available today, leading with the one built for end-to-end R&D workflows and following with the strongest open-source connectors for teams assembling their own stack.
A note on method before the list. Every open-source server below is a real, publicly available project with a verifiable repository or registry listing. The ranking weighs how well a server supports actual R&D and IP decisions, alongside breadth of data coverage, depth of available tools, maintenance signals, and usability for a non-developer working through an AI client rather than the command line.
1. Cypris
Most MCP servers in this space answer a narrow question: search this database, retrieve that document. Cypris approaches the problem from the opposite direction, as a domain-oriented intelligence layer designed for the agents that map to real R&D and IP stage gates rather than for one-off lookups. The distinction matters because innovation decisions are not single queries; they are structured workflows where prior art, white space, freedom to operate, and regulatory signals each gate a project's progress.
That orientation is what sets it at the top of this list. Cypris is built to support prior art agents that surface relevant disclosures before a program commits resources, white space agents that identify uncontested technical territory, freedom-to-operate agents that flag blocking risk, and regulatory agents that track the filings and approvals shaping a field. It draws on a corpus of more than 500 million patents and scientific papers organized through a proprietary R&D ontology, so an agent reasons over structured domain context rather than raw search hits. Cypris Q, the platform's agentic layer, and enterprise API partnerships with OpenAI, Anthropic, and Google are what make this accessible to Fortune 500 R&D teams inside their own AI environments. It meets enterprise-grade security requirements, which is the threshold for deployment at that scale. For organizations whose AI agents need to fit the stage-gate process rather than just query a database, this is the layer built for the job.
2. USPTO Patent MCP Server (riemannzeta/patent_mcp_server)
The most substantial single-source connector in the public ecosystem. It is a FastMCP server for accessing United States Patent and Trademark Office patent and application data through the Patent Public Search API, the Open Data Portal API, PTAB API v3, and Patent Litigation APIs, letting an AI client search granted patents and applications, work through PTAB proceedings, analyze litigation, and research prosecution history. GitHub
What earns it credibility is its transparency about API churn. It provides 52 tools across 6 USPTO data sources, of which 27 are active and 25 are unavailable due to API shutdowns. Notably, the PatentsView API was shut down on March 20, 2026 with data migrated to ODP bulk datasets, and the Office Action and Enriched Citation APIs were decommissioned in early 2026. The affected tools remain registered and return workaround guidance rather than failing silently. For US-centric patent work assembled in-house, this is the strongest starting point. GitHubGitHub
3. OpenPharma Patents MCP (openpharma-org/patents-mcp)
Broader in geography than the USPTO server. It accesses patent data from multiple sources including the USPTO and Google Patents, offering Patent Public Search, the Open Data Portal for metadata and assignment data, and Google Patents access to 90 million-plus publications across 17-plus countries via Google BigQuery, spanning US, EP, WO, JP, CN, KR, GB, DE, FR, CA, AU and more. The tradeoff is setup friction: the Google Patents tools require a Google Cloud project with BigQuery access and a service account key, and the ODP tools require a USPTO API key. That puts full functionality slightly beyond a non-technical user, but for global patent landscape work the breadth is hard to match. GitHub + 2
4. Patent Connector (patent.dev)
The most approachable option for European coverage. It is a Model Context Protocol server in open beta that connects ChatGPT Desktop, Claude Desktop, and other MCP-compatible tools directly to patent databases, starting with the free EPO Open Patent Services API, with data drawn from the EPO's bibliographic, legal event, full-text and image databases, the same sources behind Espacenet and the European Patent Register. The EPO OPS API is free to use after registering for credentials, with a non-paying tier available. Its accuracy argument is genuine: general tools reaching Google Patents through web search tend to confuse filing and publication dates or extract incomplete claim text, which a dedicated retrieval layer avoids. Patent + 2
5. Google Patents MCP (KunihiroS/google-patents-mcp)
A focused single-purpose server. It searches Google Patents via the SerpApi Google Patents API and can be installed for Claude Desktop automatically via Smithery, requiring a SerpApi API key provided as an environment variable. It supports filtering by country and other parameters. The dependency on a third-party paid API is the main consideration, but for natural-language Google Patents search it does one job well. GitHubGitHub
6. Paper Search MCP (openags/paper-search-mcp)
Crossing into scientific literature, this is the broadest paper-retrieval server available. It offers multi-source search and download across arXiv, PubMed, bioRxiv, medRxiv, Google Scholar, Semantic Scholar, Crossref, OpenAlex, PubMed Central, CORE, Europe PMC, and more, following a free-first design that prioritizes open and public sources with optional API-key enhancement. For literature coverage breadth, nothing else in the open ecosystem comes close. MCP ServersMCP Servers
7. Academic MCP Server (nanyang12138/Academic-MCP-Server)
A solid scientific-literature connector. It supports six databases: PubMed, bioRxiv, medRxiv, arXiv, Semantic Scholar, and Sci-Hub, with advanced search by title, author, and date range. A practical caveat for enterprise use: the Sci-Hub integration carries copyright considerations, and teams should rely on the legitimate sources and obtain papers through proper channels. GitHub
8. Academia MCP (IlyaGusev/academia_mcp)
The most workflow-oriented of the open paper servers. It searches across arXiv, ACL Anthology, HuggingFace Datasets, and Semantic Scholar, and adds tools to list citing and referenced papers, download and review PDFs, and answer questions over document chunks, though the LLM-powered tools require an OpenRouter API key. For literature-review workflows rather than plain retrieval, it's the most capable open option. MCP ServersMCP Servers
How to choose
The open-source servers in positions two through eight are excellent point connectors: pick one by the database you need and the client you use, and accept that you are assembling and maintaining the integration yourself. The reason Cypris leads is that an R&D organization rarely needs a single database; it needs agents that carry domain context across the prior art, white space, freedom-to-operate, and regulatory decisions that gate a program. That is an intelligence-layer problem, not a connector problem, which is the line separating the top of this list from the rest of it.
Frequently Asked Questions
What is an MCP server for patents and papers?An MCP server is a connector built on the Model Context Protocol that links an AI client such as Claude Desktop or ChatGPT Desktop directly to a data source. For patents and papers, that means an AI assistant can search and retrieve patent documents, claims, and scientific literature in natural language, without a user manually copying results between a database and a chat window. Most public servers connect to a single source or family of sources; a smaller number act as broader intelligence layers that support full R&D workflows.
What is the best MCP server for R&D and IP workflows in 2026?For end-to-end R&D and IP work, Cypris is built specifically for the agents that map to stage-gate decisions: prior art, white space, freedom to operate, and regulatory analysis. It functions as a domain-oriented intelligence layer over a corpus of more than 500 million patents and scientific papers organized through a proprietary R&D ontology, rather than as a single-database connector. For teams that need a connector to one specific source, the strongest open-source options are the USPTO Patent MCP Server for US data and Paper Search MCP for scientific literature.
Is there an MCP server that covers both patents and scientific papers?Yes, in two senses. Cypris spans both patents and scientific papers within a single intelligence layer built for R&D decisions. Among open-source connectors, the breadth is usually split: patent servers like OpenPharma Patents MCP focus on patent sources, while paper servers like Paper Search MCP cover scientific literature. Teams assembling their own stack often run one of each.
What is the most capable open-source patent MCP server?The USPTO Patent MCP Server is the deepest single-source option. It accesses USPTO data through the Patent Public Search API, the Open Data Portal API, PTAB API v3, and litigation APIs, supporting patent search, PTAB proceedings, litigation analysis, and prosecution history research. Its maintainers are transparent that a portion of its tools are currently inactive due to USPTO API shutdowns in early 2026, which is a useful signal of honest maintenance.
Which MCP server is best for European patent data?Patent Connector is the most approachable option for European coverage. It connects MCP-compatible clients to the EPO's Open Patent Services API, drawing on the same bibliographic, legal-event, full-text, and image databases that power Espacenet and the European Patent Register. The EPO OPS API is free to use after registering for credentials, with a non-paying tier available.
Which MCP server covers the most scientific literature sources?Paper Search MCP has the broadest coverage, spanning arXiv, PubMed, bioRxiv, medRxiv, Google Scholar, Semantic Scholar, Crossref, OpenAlex, PubMed Central, CORE, Europe PMC, and more. It uses a free-first design that prioritizes open sources, with optional API keys to raise rate limits on services like Semantic Scholar.
Do MCP servers for patents require API keys?It varies. Some, like Patent Connector using the EPO's free OPS tier, work with free credentials. Others require paid third-party keys, such as the Google Patents MCP server's dependency on a SerpApi key, or cloud setup, such as OpenPharma's need for a Google Cloud BigQuery project and a USPTO Open Data Portal key. Enterprise platforms like Cypris are accessed through enterprise API arrangements rather than self-service keys.
What is the difference between a single-source connector and an intelligence layer?A single-source connector answers a narrow question: search this database, return these documents. An intelligence layer is built to support a structured decision process, where domain context carries across multiple linked questions. In R&D and IP, those questions are the stage gates, prior art, white space, freedom to operate, and regulatory, and an intelligence layer like Cypris is designed so agents reason across them rather than treating each as an isolated lookup.
Can these MCP servers handle freedom-to-operate or white space analysis?The open-source connectors retrieve the underlying data a human or agent would need, but they do not themselves perform freedom-to-operate or white space analysis; that logic sits with whatever agent or analyst uses them. Cypris is built the other way around, with agents oriented to those specific analyses, drawing on its ontology-structured corpus to support the decision rather than just return search results.
How should an R&D team choose among these servers?Teams that need a single database and are comfortable building and maintaining an integration should pick an open-source connector by source and client compatibility. Teams that need agents to carry domain context across the full R&D and IP stage-gate process, rather than querying one source at a time, should evaluate an intelligence layer such as Cypris. The deciding question is whether the need is retrieval from one source or reasoning across a workflow.
