My goal was to reach 1% of Google's index. This is 4 billion pages across 33.5 million domains. I was inspired by the earlier attempts shared here: "Building a web search engine from scratch with 3B neural embeddings" [1] and "Crawling a billion web pages in just over 24 hours" [2]. I decided not to replicate them and instead used CommonCrawl corpus, built a pipeline, and threw the results into a classic full-text search index.
First, I got my hands on a dedicated Hetzner server with 4x10TB HDDs. It allowed me to crunch CommonCrawl data and stay within my hobby budget. The backend is served from a similar "home-grade" server but with NVMe disks instead. The compressed index weighs just over 2TB because I cap content length at 4KB per page (p50=2.9KB, p99=46KB) to fit it on disk. I picked Rust for my project. After years of big data engineering on the JVM, Rust felt like a breath of fresh air. I ended up not using any special data processing libraries - the pipeline is relatively straightforward and runs without crashes. As for full-text search, tantivy was an obvious choice. It took some time to get single query latency down to an acceptable ~350ms, although concurrent requests quickly max out disk t/p at 6.1GB/s and latency starts to climb. The search algorithm is primitive by modern standards: top 1,000 candidates are retrieved using BM25 (text relevance signal) and then reranked with PageRank [3].
Turns out you can fit the whole internet in a single server rack. The challenge here, of course, is to keep it up to date, especially if you don't have unlimited (crawl) budget. Right now I'm looking for funding or sponsorship to get extra hardware. That would allow me to scale the project and keep it free to use. I can also periodically dump the index on HuggingFace for research purposes. Any suggestions are welcome.
[1] https://news.ycombinator.com/item?id=44878151