
From Raw Patent Text to Bag-of-Words Features
Feature extraction from patent text is the process of automatically pulling out the lexical and semantic building blocks of a document so that algorithms can work with it, and it underpins later stages such as query expansion, document classification, and topic modelling. Before any AI system can classify documents, expand a search query, or measure similarity between patents, it needs a way to represent the text itself in a form a machine can handle. The research team approached this in three layers, each capturing a different level of linguistic detail: single words, multi-word terms, and finally word embeddings that capture meaning itself.
At the most basic level, the system's document representation relies on the "bag-of-words" model. Each document is treated as a simple collection, or "bag," of the words it contains, with no regard for grammar, sentence structure, or word order. On the surface, this might sound like a crude simplification, and in some ways it is. Yet this representation has proven remarkably successful in information retrieval and document classification, largely because the sheer number of words involved makes it straightforward to quantify how locally relevant a given word is. The standard measure for this is term frequency-inverse document frequency, better known as TF-IDF, a long-established technique that traces back to foundational work by Spärck Jones (1972) and Salton and McGill (1986).
Normalising Words So Variants Match
Individual words only become useful features if the system recognises when different surface forms are really the same word. To handle this, the researchers applied basic linguistic pre-processing that neutralises insignificant orthographic differences between words that mean the same thing:
- Letter casing, so "Bayesian learning" and "bayesian learning" are treated as one term
- Non-ASCII characters, so "naïve Bayes" and "naive Bayes" match
- Spelling variations between English dialects, such as "nearest neighbour" and "nearest neighbor"
- Spelling mistakes, which are corrected automatically
Beyond this, the system applied lemmatisation and stemming, techniques that strip words down to their underlying root or core meaning. For example, "transportation," "transported," and "transporter" all map to the shared root "transport," so the system can recognise that these different forms are about the same underlying concept. This kind of normalisation is a basic building block of any serious patent information search, since patent authors rarely use identical wording for the same idea.
Multi-Word Features and the Limits of N-grams
Single words alone miss a great deal of meaning that only emerges when words appear together in specific combinations. To capture relationships between individual words, the researchers explored two complementary approaches built around word co-occurrence: fixed-size text windows known as n-grams, and a separate method focused on domain-specific multi-word terms.
N-grams preserve the local context around individual words, so they can be layered on top of the bag-of-words representation to allow a more fine-grained comparison between documents. But they have a real limitation. They divide text into fixed-size blocks based purely on physical proximity, with no regard for the logical, syntactic, or semantic relationships between words, and this can cause real information to be lost. The researchers illustrate the problem with an instructive example. One document contains the phrase "the way of doing things on the Internet has evolved," while another contains "five ways the Internet of Things is transforming businesses." Represented as bags of words, the two look quite similar, since both mention "Internet" and "things." But only the second refers to the Internet of Things, the standard technical term for the interconnection, via the internet, of everyday computing devices that can send and receive data. A bi-gram approach, which pairs up two consecutive words, would fail to capture this distinction. Tri-grams would successfully capture "Internet of Things" as a single feature, but even that is not enough for longer technical terms such as "Internet Small Computer System Interface." A more flexible approach was clearly needed, one capable of capturing important phrases regardless of how many words they span.
Extracting Technical Terms With FlexiTerm
This is where multi-word terms come in. They act as the linguistic representation of domain-specific concepts, and because they convey technical and scientific information as coherent units, they tend to get lost when text is chopped mechanically into fixed-size n-grams. To handle them properly, the research team used a locally developed software tool called FlexiTerm, built to extract multi-word terms from text on the fly and drawing on prior research by Spasić and colleagues (2013, 2018). One especially useful capability is its ability to link acronyms to the full terms they stand for, for instance connecting "Internet of Things" to "IoT," or "Internet Small Computer System Interface" to "ISCSI." Unpacking an acronym into its full form makes its component words searchable, turning an opaque abbreviation into a usable feature for further document analysis. If you want to see how an AI-powered tool copes with dense technical terminology in practice, you can try a free novelty search with WOIPS on your own invention.
The system was also designed to group different variants of the same underlying term into a single feature. One documented example is simple orthographic variation, "bottom hole assembly" versus "bottomhole assembly," both of which the system links to the shared acronym "BHA." A more complex example involves genuine syntactic variation, where the word order itself differs: "network functions virtualization" versus "virtual network function." These produce two different acronyms, NFV and VNF, for what is fundamentally the same concept. Even here, stemming recognised that "virtual" and "virtualization" share the same root, connecting the variants through their core meaning rather than their surface form. To make the extracted terms browsable, the system could also organise them into dendrograms, tree-like diagrams that group related terms together by type.
Word Embeddings: Capturing Meaning Itself
The single-word and multi-word approaches both represent words as discrete, separate variables, which means there is no built-in way to compare how similar two words are in meaning or to capture subtler semantic relationships between them. To close this gap, the researchers turned to a fundamentally different kind of representation: word embeddings.
Word embeddings build on distributional semantics, a bottom-up approach to meaning centred on the distributional hypothesis, which holds that words appearing in similar contexts across a body of text tend to have similar meanings. In practice, each word is represented as a real-valued vector of relatively low dimensionality, learned automatically from text using techniques such as neural networks or dimensionality reduction. By generalising the typical context in which a word appears, embeddings preserve meaningful relationships between words, expressed geometrically as distances and directions within the vector space. They also help sidestep the "curse of dimensionality," a phenomenon first described by Hughes (1968), in which the performance of many machine learning algorithms degrades as the number of dimensions grows too large.
Why Domain-Specific Embeddings Matter
For this study, the researchers used word2vec (Mikolov et al., 2013), a well-established, state-of-the-art embedding algorithm, to train separate word embeddings for each of the three chosen domains. The result was genuinely domain-specific word representations rather than one-size-fits-all embeddings, and this matters a great deal in practice. Take the word "driver," documented in one of the report's appendices. It can refer to a physical mechanical object in civil engineering, a piece of software in computer technology, or a person operating a vehicle in the transport domain. Because the embeddings were trained separately for each domain, they captured these different meanings through the word's distinct relationships to other similar or related words.
In the transport domain, "driver" sits close to synonyms such as "vehicle-operator," more specific terms (hyponyms) such as "cyclist," and related concepts such as "passenger." In computer technology, the very same word clusters instead near synonyms such as "controller" and related terms such as "I/O," reflecting its entirely different technical meaning. This ability to disambiguate meaning from domain context is a real step up from the simpler single-word and multi-word approaches. It also shows why a reliable prior art search depends on understanding context rather than matching keywords alone. To see how an AI-powered search handles your own invention, explore the WOIPS Novelty Search service.
Frequently Asked Questions
What is feature extraction in patent search?
Feature extraction is the process of automatically pulling lexical and semantic building blocks, such as single words, multi-word terms, and word embeddings, out of raw patent text so that algorithms can classify documents, expand queries, and compare patents.
What is the bag-of-words model and why is it used for patents?
The bag-of-words model represents each document as a collection of the words it contains, ignoring grammar and word order. It works well for retrieval and classification because measures such as TF-IDF can quantify how relevant each word is to a document.
Why are n-grams not enough for technical patent terminology?
N-grams split text into fixed-size blocks based on proximity alone, so they can miss longer technical terms such as "Internet Small Computer System Interface" and fail to separate phrases like "Internet of Things" from unrelated mentions of "Internet" and "things."
What does FlexiTerm do in a patent search system?
FlexiTerm extracts multi-word technical terms from text, groups spelling and word-order variants of the same term, and links acronyms to their full forms, which makes the individual words inside those terms searchable.
Why train word embeddings separately for each technology domain?
The same word can mean very different things in different fields. Training embeddings per domain lets the system place a word like "driver" near vehicle-related terms in transport, but near terms such as "controller" in computer technology.
This article is adapted from: UK Intellectual Property Office, AI-assisted patent prior art searching – feasibility study, April 2020, ISBN 978-1-910790-80-9, © Crown Copyright 2020, licensed under the Open Government Licence v3.0.
