To Be Indexed Is to Exist: Between Relevance, Legitimacy, and Web Archives

According to the online documentation provided by Google for webmasters:

"TheGoogle indexis like a library catalog, which provides information about all the books available in the library."
— Index pages to include in search results, Support.google.com

Google’s job is therefore clear. It involves scanning all the content published on the Web and classifying the information. The classification method takes the form of a search engine that returns results based on certain criteria (which may or may not be explicitly defined).
Library users and librarians are familiar with the Dewey Decimal Classification system—it’s just one system among many! The key is to find your way around the aisles and shelves. When it comes to the Web, if you imagine the largest and most complex library, it’s as if, with every online publication, someone walked into the room to place a book or a single page. Then you have to figure out what to do with it. And that happens billions of times a day. 

A Bet

When we talk about SEO, we naturally tend to think about how to gain visibility. The words “ranking” or “position” are probably the first to come to mind. However, achieving a high ranking is already an advanced step. After all, nothing is possible if indexing isn’t effective.

Aside from that, one has to appreciate the technical feat—and even the gamble—involved in wanting to

  • browse all online content to store information,

  • analyze the topics,

  • process the data as a whole in order to produce a relevant ranking.

But Google wouldn't be Google if all those bots weren't constantly scouring the Web to collect all that data, nonstop and with the goal of being exhaustive.

What's the difference between Google and a library?

What would you say if I asked you to name an encyclopedia? 

Wikipedia?

  • Indexing by Index Cards

  • The very existence of an article is up for debate. What justifies the absence of real, recognized professionals in their fields, while fictional characters like Pikachu have their own Wikipedia pages?

  • Issues with primary sources. A source must have been cited numerous times by other entries. This is an interesting method for selecting topics to cover, but it can have its flaws.

Larousse Encyclopedia

  • Academic legitimacy.

Dictionary

Every year, we see the debate over which new words are accepted into the dictionary. This is a sign of how the French language is evolving, but inclusion in the dictionary doesn’t happen overnight either. It’s easy to imagine the debates that must take place at the Académie Française.

When it comes to Google, anyone can post content online. That is precisely what makes the web so rich and what has enabled the emergenceof major collective initiatives. Online publishing is a means of communication. By creating content, we bring a topic into the index of the search engine that has become the leader in information publishing—without borders or limits. But this ability to publish without mandatory academic validation also creates complexity: the question of legitimacy is important.

While the works in a library have been vetted at some point (by a publisher, a librarian, etc.), online content indexed by Google is, by its very nature,unmoderated.Indeed, in the case of Wikipedia or Larousse, it is people who collectively approve the publication of a definition. Google, on the other hand, offers an automated, algorithmic ranking system.There are certainly times when manual intervention occurs, but that’s not always a good sign (#penalty). 

If we draw a parallel with the Bibliothèque Nationale de France, it has its own robot, known as a “crawler.” It even has a first name:Heritrix.  Although WebArchives may seem similar, WebArchives’ methods for its Wayback Machine are undoubtedly different. The challenges and objectives are different as well. As far as I’m concerned, I see theWayback Machine more as a form of historical versioning of the Web, whereas the BnF has taken on a role of collection. It’s even a responsibility.

The Web is like Santa Claus—it never forgets

Let’s keep in mind that the procedure for exercising the right to be forgotten involves having information de-indexed, not deleted. It is not a complete deletion, even though it is widely acknowledged that content that cannot be accessed is difficult to access. It’s the difference between tearing down a house and erasing the words on the signs that point the way to it. If, on top of that, you destroy the roads by removing the links to that page, it will be even more difficult to access the content in question.

Relevance or Legitimacy

BetweenGoogle E-A-Tandcore updates, relevance is always the key issue. Content that is indexed and ranked by a search engine is, in principle, destined to endure over time. What if we assumed that we are creating tomorrow’s archives every day? What should we leave for posterity? This also raises questions of ownership and digital pollution.

It is therefore worth remembering that the concept of legitimacy is relative. Content that may seem unimportant, useless, or even ridiculous to you might be exactly the opposite to someone else. One might think that conversations on Twitter—for example—are best discarded. Yet the suspension of Donald Trump’s personal account also raises the question of the archives of a former U.S. president’s public statements. Admittedly, this deletion is justifiable for many reasons, but shouldn’t we have access, in one way or another, to the words of someone who has influenced the history of a country—if not the world? Moreover, the buzz on Twitter is a goldmine for gauging public opinion; otherwise, tools likeVisibrainwouldn’t exist. 

Before we rush to delete everything and pass judgment on whether certain information online has a right to exist, let's wait and see. We might be surprised by the content that will outlive us.

___________

This article is based on my notes from the talk I gave at the Google Search Central Meetup in Paris (October 13, 2022). This Meetup, organized by the Google Search Liaison team (Zurich), took place on November 13, 2022, at Google’s offices in Paris on Rue de Londres. I’d like to extend my warmest thanks to Martin Splitt, Myriam Jessier, and Aymen Loukil for organizing this event. Thanks also to Rebecca Berbel for being a fantastic moderator. And of course, thank you to all the participants!

Rebecca Berbel and Syphaïwong Bay at the Google Search Central MeetUp in Paris on October 13, 2022.

The presentation did not include slides; it was delivered without visual aids. It is an expanded version ofone of my articles written for Reacteur (Abondance).

Previous
Previous

Storytelling, Copywriting & SEO: The Startup Package for Naming and Promoting Your Innovation in the Market

Next
Next

How can you enhance your content using SEO- and UX-friendly methods?