wikipedia – Byte64

Visualizing Retrieval

May 7, 2026

—

by

Andrew

Yesterday I created a Byte64 linkedin page, so clearly had to create some visualizations to use as a background image. No AI slop would do. The first thing I tried was visualizing the PFORs the encode my search index. I don’t have a good sense of what the spacing is like between local document ids…

Categories and Suggestions

Apr 28, 2026

—

by

Andrew

in Algorithms, Data Processing, Data Structures, Search

Andrew

Recently I’ve been trying, somewhat unsuccessfully, to find wikipedia search filters that don’t fit into my retrieval model. I was hoping to find some natural low cardinality, high coverage fields that people could use as filters in their queries. Imagine you’re on a product page for Bombas socks. You’ll find filters for product type (Men/Women/Kids/Sport),…

PageRank on Wikipedia

Apr 17, 2026

—

by

Andrew

in Algorithms, Data Processing, Search

Andrew

PageRank is a famous scoring function invented and deployed by the Google guys in early days of websearch. It assigns each webpage a score based on the scores of webpages that link to it. As you can see, it’s a recursive definition, but if you use the right formula, then it’ll converge to something meaningful.…

BM25 & Search Index Encoding

Apr 2, 2026

—

by

Andrew

in Data Structures, Search

Andrew

Okapi BM25 is a standard ranking formula that has been used in search engines since the 1980s. For each word in a query, it uses the frequency of that word in a document, the length of the document and the number documents that contain the word to decide how significant the word is for the…

Vector Embeddings

Mar 26, 2026

—

by

Andrew

in Data Structures, Search

Andrew

Search engines make use of AI to improve their search results. There are AI models that can understand the meaning of a sentence or document. They often present their results as embeddings of the document space into a vector space: D→ℝnD \to \mathbb{R}^n. These are called vector embeddings. Once you’ve found the embeddings for your…

Generating User Data

Mar 18, 2026

—

by

Andrew

in Coding with AI, Search

Andrew

A search engine depends on a feedback loop of users making queries, following links, returning to the search page and rewriting their queries. All these feed into an understanding of whether they’re finding the results they’re looking for. Bootstrapping a system like this is difficult because you don’t start with any users and your search…

Search Indexes and Memory

Mar 12, 2026

—

by

Andrew

in Data Structures, Search

Andrew

I’ve been working on v0 of the search engine which requires building a search index in the form of “posting lists”. I’ve build a pipeline that reads wikipedia documents, outputs all the words it finds and emits key-value pairs of (word, url). Then we group by word so we can lookup all documents that word…

Writing SSTables with Beam

Mar 10, 2026

—

by

Andrew

in Coding with AI, Data Processing, Data Structures

Andrew

Apache Beam is an open source system for processing large datasets. It has both a realtime and a batch processing mode. The batch processing mode is based on Google’s internal Flume framework which I had the pleasure of using for 7 years while processing Android telemetry. It’s also the perfect system for building a search…

SSTable on SSD

Mar 5, 2026

—

by

Andrew

in Google Cloud

Andrew

When I first put the SSTable server on Google’s Compute Engine, I opted for the cheapest machine: an e2-small (2 cores and 2GB ram) with a standard persistent disk of 20 GiB. When I ran the server, each request was taking hundreds of milliseconds. Installed iotop and found I was getting at most 8 MiB/s…

Building an SSTable

Feb 28, 2026

—

by

Andrew

in Coding with AI, Data Structures

Andrew

SSTables are a critical piece of technology that holds up the modern web. It’s the basis for most modern databases, search backends and many other technologies. What it provides is a reasonably fast way to perform lookups in large datasets. SSTables are sorted string tables meaning both our lookup keys and the resulting values are…

Tag: wikipedia