BERT was announced at the end of 2018. A year later, the update is live in English-speaking countries — but what impact does it have, and what changes need to be taken into account?
Most of the time, internet users go to Google to look for information they do not have the answer to; they are looking to inform themselves. Yet it can sometimes be difficult for Google to find the right answer to a query when the user does not know how to phrase it clearly.
Over the years, Google has improved a great deal in trying to understand what information users are looking for when they type keywords. Now it wants to get better at what are known as natural language, or conversational, queries thanks to machine learning (a process in which the machine learns by itself, using the data it collects).
What exactly is natural language?
Today, with the arrival of connected devices that make voice search possible, internet users are using spoken language: they address the machine as they would in a real conversation with another person. Until now, however, what Google understood best were queries made up of one or more keywords.
To illustrate what natural language is, we can use the example of the question: “What can you do when some people are bored?” → We are looking for an answer that would give us an idea of how to remedy that problem. However, the results displayed by Google may not be relevant to the intent behind the search. We may not find the answer to our question in the items offered by Google’s SERP, because the substance of the question has not been understood by the search engine.
In this query, the search engine only analysed the keyword “bored” without understanding the context of the question that would have helped it display the right results.
That is why Google wants to improve on this point. It wants to understand what the user is looking for when they express their query naturally, in spoken language. So it has introduced the BERT algorithm.
Bidirectional Encoder Representations from Transformers, alias BERT
BERT is a significant update to the Google Search algorithm. It is the artificial neural network created by Google with the aim of better understanding the user’s natural language. It concerns long queries — those that contain a lot of words and resemble a sentence.
Before the update, Google relied on the keywords typed during a search in order to match them with similar content. Today, with the BERT algorithm, it is all of the words that carry meaning within the content.
BERT aims to understand precisely the relationship between the different words in a sentence, their position, their importance and what they mean, in order to obtain ever more accurate results. To do that, it has to analyse the query as a whole and no longer simply take keywords into account individually.
Let’s take an example with the query “Can I book a hotel for my friend?” → Google analysed the keywords “book” and “hotel” but did not understand that “for my friend” was also an important element in the question.
With the new BERT algorithm, the query will be better analysed in order to offer articles that answer the question asked more precisely.
However, for the time being this update will only affect one query in ten — but Google hopes that we will subsequently get into the habit of no longer searching only with keywords and that we will gradually turn towards broad natural language queries.
How BERT works
The BERT algorithm uses two strategies to understand the meaning of queries:
1. Masked language
BERT tries to predict the rest of the query. This is a method that consists of masking certain words at random with what is called a [MASK] token in order to predict them (generally around 15% of the query). The algorithm has to analyse the sentence as a whole in order to guess the words that might be hidden behind each token.
For example: with the query “Buy the book Les Misérables by Victor Hugo”, the algorithm will see “Buy the [MASK] Les Misérables by [MASK] Hugo”. → It will analyse the query in order to anticipate what might lie behind the masks.
2. Next sentence prediction
BERT analyses sentences bidirectionally — in other words, it is able to analyse and judge whether two sentences are related or not.
For example, if we take a sentence 1 and a sentence 2:
Sentence 1: my computer won’t connect to the internet. Sentence 2: we can’t start the video conference. → The two sentences are related.
Sentence 1: my computer won’t connect to the internet. Sentence 2: my cup of coffee is ready. → The two sentences are not related.
The algorithm classifies the two sentences and calculates the probability of their being connected to each other, in order to better understand the context of the search.
Conclusion
BERT will only promote quality content, so you need to refine your content carefully in order to be as precise and relevant as possible, because it will naturally no longer favour sites with neglected content.
For the time being the algorithm is only available in English-speaking countries, because the algorithm’s artificial intelligence has the ability to teach other languages to reproduce the same pattern. The first tests therefore began in the United States and will arrive in France very soon!
BERT is not a fixed project — it will be constantly developed — but as things stand it already represents a major step forward.
Looking ahead, another advance is already on the schedule: the ALBERT algorithm, which should succeed BERT by offering an improvement in terms of its size and weight. Watch this space!
If you enjoyed this article, take a look around our blog for more reading!