Has Google search been neutered? ================================ Author: Michail Konstantinos Dimopoulos Date: 2023-11-23 Site: https://mcdim.xyz Note: this article was written in 2023 and was never published. I'm publishing it now in 2026 for archival, and because I think it has retrospectively become more interesting with recent developments. I remember the days when using google meant you could find almost anything you wanted, without even using advanced syntax. You'd find links for forum discussions, quality articles, personal blogs & projects. It could actually be described as surfing. It's not like you cannot find these things at all nowadays, but I can't be the only who has noticed that, for years now, searching on google has gradually become a worse exeprience. You know it's really bad when even redditors are catching on to it. About 6 gigatrillion results (0.0014 nanoseconds) ------------------------------------------------- It is relatively well known that the report of how many results google has retrieved with your query is off. It's not uncommon for me to be told my search has millions of results, only for them to not last more than 2 pages. There's nothing wrong with the number not being very accurate, it's an approximation afterall. It would be unnecessarily inefficient to trasverse all of the data in their database just to give you an accurate number. "Approximation", however, implies some kind of approximation. Claiming you have hundreds of thousands or even millions of results when, in reality, you only have 141 isn't an approximation. It's a lie. I just did a little quick & dirty experiment. I searched a relatively non-controversial/political technical query that I knew wouldn't have too many or too few results: "unifi" "openvpn" "expected". Google reported around 62.500 results. I counted them, I got less than 200. I even enabled the option to get duplicate/low-quality results. Before that, it was around 50. Who knows? Perhaps Google's approximation tolerance range is about 100%! This has always been an issue on the site, but it may be the worst it's ever been, and I don't understand why google still has this non-feature. The number is not even close half the time, then what's the point of even reporting a number? I can't shake the feeling that it's there to give you a false perception of freedom & variety. Google knows most people won't go past the 3rd page (are pages on Google even a thing anymore?), and that they won't realise that that is where it ends. It's like a walled garden where the walls are transparent. You only realise there are walls when you try to go past them. This is particularly eerie when you realise that the situation is worse with controversial political topics. I remember during the pandemic, searching for the vaccines you would get the same CNN article (or variations of it) over and over again for 3 pages and then nothing. I decided to do a second experiment, I searched for covid vaccine side effects, for which Google reported 481.000.000 results for 0.65 seconds. Oh wow! Thanks Google for providing me with all these pages & diverse opinions on the topic! I scrolled all the way to the bottom of the results, which stopped at 173. Google didn't consider any other pages high quality enough to be included in the search. In the entire English-speaking world, only 173 articles on the vaccine side effects were high quality enough. I mean how many articles did you expect the Google propaganda ministry to manually review all by themselves? Cut them some slack! I clicked on the option to allow duplicate & low quality results, scrolled all the way to the bottom, copied everything and pasted it to a file. I then greped that file for links & piped that to wc -l, which returned the true number of results: 434. These are the 434 webpages you're allowed to read about covid vaccine side effects. Consider yourself informed. We've had AI search all along! ------------------------------ If you've been wondering why articles on your google results are entirely unhelpful, use corporate-y language & read as if they were written by Chat-GPT, that's because they most likely were. The epidemic of trash articles written by AI is more real than COVID. What's worse is that these soulless websites have not only maxxed out their SEO, but have reached ad-to-content ratios previously thought impossible. Many people make an effort to avoid these irritating websites by appending site:reddit.com or something along those lines to their search queries, with the goal to find actual human discussions. While that can work, I've found reddit to be just as frustrating & even more insidious in its political censorship. Plus, whether accounts on reddit are actual humans interacting is rather debatable at this point. Most discussions on large subreddits are sanitized, and besides, these redditors are the primary way Chat-GPT and other text-spitting neural networks are trained, so expect to find similar cookie-cutter opinions. You could use these to be avoid reddit, or at least be more website-agnostic: inurl:forum, inurl:thread, intext:reply I was going to give more detailed information on how to effectively use google's advanced search syntax to avoid trash AI sites, however, I don't consider these to be solutions, and Google doesn't seem to like them either: Using advanced search syntax? Not in this site, punk! ----------------------------------------------------- The main selling point of google to hackers, and me as well, is the powerful list of commands you can give google to perform precise searches. These include things like -keyword, "keyword", AROUND(0), filetype:pdf, site:example.com etc. In fact there is a whole science in the field of cracking that involves finding private information through web searching with google's advanced syntax, called [1]Google Dorks. Traditionally, these were used by powerusers, but nowadays, these are necessary even the most basic of web searching, at least if you don't wish to be bombarded with spam AI articles. Google, however, doesn't seem to think so, as they've been making it more and more frustrating to use them by spamming you with "I'm not a robot" captchas. And they exponentially increase the more you use them. Of course this is caused by many crackers using them, but as a result they're making the only way to properly use the engine a not-that-great experience. Absit omen a man wants to read forum threads on how to fix his car, rather than read auto-generated junk posts filled with megabytes and megabytes of ads and trackers. Literally 1984 -------------- No arguments needed. DuckDuckGo and other "private" search engines are not solutions Well DuckDuckGo definitely is not, as its very founder & CEO, Gabriel Weinberg, has admitted on twitter that [2]DDG implements political censorship. Apparently now they also include a fact-checker! Thanks DuckDuckGo for having my back against evil misinformation! But, just like other private search engines, it was never going to be a solution, anyway. Its model is the same as that of Google: a centralized proprietary technology, only with the added marketing towards privacy enthusiasts. In reality, you can never know what DDG, or any other private search engine, actually does with your search queries & data. Their claims are unverifiable & you can only trust their word. Sure, it may not be as visibly intrusive as Google is, but once its on their servers, its theirs and we can never know what they do with that data, unless there's a server breach proving that they store them. Similarly, you can simply never know whether the search engine you use is censoring websites or not. All you can do is trust their word. It just so happened that Gabriel Weinberg is an absolute fool and admitted to it. If he didn't, we simply couldn't have known. SearXNG, wiby or similar search engines can be a solution [3]SearXNG (a maintained fork of SearX), unlike most search engines, is a free software meta-search engine that can be self hosted. It retrieves search results from other search engines, acting as a proxy & is very customizable. Likewise, [4]wiby.me is free software & can be self hosted. It has its own internal database with sites authored & submitted by actual people and, as such, has essentially no spammy AI articles. If you want more information on alternative search engines, there's this [5]great article at digdeeper.club that analyzes them in detail. Do we even need search engines? ------------------------------- Search engines are seen as an integral part of the internet, if not as the internet itself. In reality, the web can & has functioned without search engines. "Surfing" and discoverability can happen through word of mouth, links on personal sites & blogs, webpins, IRC/Matrix/XMPP, E-mail, RSS/Atom, and webrings like [6]the one I belong to. I, and others, have [7]link pages on our websites to promote sites we find interesting. I've found a lot more intereting things this way than ever using a search engine, except perhaps for [8]wiby.me. That way we don't have to worry about being tracked by a mediator, or spammed by AI garbage articles, as everything is curated by actual humans. Actually searching for something specific over the whole web, however, is a service that only search engines can provide, and for this reason we're stuck with them to some degree. However, as the Googles of the internet are gradually ceasing to be proper search engines, alternatives will have to take their place. We just have to make sure it's the right ones that do. References 1. https://en.wikipedia.org/wiki/Google_hacking 2. https://digdeeper.club/articles/search.xhtml#ddg 3. https://github.com/searxng/searxng 4. https://wiby.me/ 5. https://digdeeper.club/articles/search.xhtml 6. https://heaventree.xyz/ 7. https://mcdim.xyz/links.html 8. https://wiby.me/