# Searching your screen by meaning, on the Mac.

In Earlyn's test of 12 search-by-meaning queries in English and Turkish, EmbeddingGemma 300M and bge-m3 ranked the right document first all 12 times, Apple's NLContextualEmbedding 4 times, and a converted multilingual-e5-small model once.

| Model | Right document ranked first | Notes |
| --- | --- | --- |
| EmbeddingGemma 300M (Q8) | **12 / 12** | 334 MB; chosen for Earlyn |
| bge-m3 | 12 / 12 | Much larger download |
| Apple NLContextualEmbedding (mean-pooled) | 4 / 12 | Built into macOS |
| multilingual-e5-small (a GGUF conversion) | 1 / 12 | The conversion we tried, likely not the model itself |

## What made the difference: the window title

Screen text is mostly menus, buttons and sidebars. Embedding it alone ranked badly. Putting the window title in front of the first 1,500 characters, in EmbeddingGemma's `title: … | text: …` layout, fixed most failures: “müzik dinlemek” (listening to music) found the YouTube tab, and “rakip ürünün geliri” (the competitor's revenue) found a search for a rival product's monthly revenue.

## How it runs in Earlyn

- Each new moment gets a vector after it is stored; older ones are filled in at background priority, 32 at a time.
- Vectors are stored as 8-bit integers next to the encrypted text; an embedding takes 25–70 ms.
- A search compares the question with every stored vector using Apple's Accelerate framework, keeps results close to the best one, and shows one moment per window.
- Matches by meaning appear in search marked “related”, next to the matches by word.

> **Caveat** Twelve queries from one person's own history is a sanity check, not a benchmark. It was enough to rule out the weakest options for this use.

---
Canonical: https://www.earlyn.app/research/local-semantic-search · Updated 2026-10-02
