Embedding a few hundred starred repos means a few hundred README fetches and a few hundred model runs. That takes minutes, can hit GitHub's rate limit halfway through, and shouldn't die if the user navigates away or quits the app. So the HTTP request only kicks things off, and a background worker does the rest:
1POST /enrich/starred/run
2 |
3 +--> worker.start()
4 |
5 +--> enqueueAllStarredRepos() GitHub GraphQL, one page of stars at a time
6 |
7 v
8 +------------------+
9 | queue (SQLite) | one job per repo, deduped by repo id
10 +------------------+
11 |
12 v batches of 10
13 Worker: fetch README -> build text -> EmbeddingGemma -> upsert into PGlite
14 |
15 +--> patchEmbedActivity() --> pub/sub --> SSE (chapter 4)
We use Conveyor, a job queue with a BullMQ-style API (Queue, Worker, retries with backoff, deduplication, pause/resume, batch processing) and pluggable stores. We use the SQLite store (@conveyor/store-sqlite-node), so the queue is just a file under ~/.config/tangerine-desktop/queues/.
That fits the app: one Deno process on one machine. There's no Redis to run, jobs survive restarts because they're on disk, and the whole engine lives in the same process as the UI server and the model.
Each queue gets its own SQLite file (src/lib/worker/store.ts):
1export function createWorkerStore(options: { name: string }) {
2 const filename = resolveLocalPath(
3 process.env.QUEUE_DATABASE_PATH ?? process.env.QUEUE_DATABASE_URL,
4 `queues/${options.name}.db`,
5 );
6 mkdirSync(dirname(filename), { recursive: true });
7 return new SqliteStore({ filename });
8}
1
2export const starredRepoEmbedStore = createWorkerStore({ name: "starred-repo-embed" });
3await starredRepoEmbedStore.connect();
One job per repo. The payload is just what we already got from the stars list, so the worker doesn't need to re-fetch metadata (queue.ts):
1export interface EmbedRepoShape {
2 id: string;
3 owner: string;
4 name: string;
5 description: string | null;
6 languages: string[];
7 tags: string[];
8}
9
10export const starredRepoEmbedQueue = new Queue<StarredRepoEmbedJob>("repo-embed", {
11 store: starredRepoEmbedStore,
12});
13
14export const enqueueStarredRepoEmbedJobs = async (jobs: StarredRepoEmbedJob[]) => {
15 const created = await starredRepoEmbedQueue.addBulk(
16 jobs.map((job) => ({
17 name: "embed-repo",
18 data: job,
19 opts: {
20 deduplication: { key: `repo-embed:${job.id}` },
21 attempts: 5,
22 backoff: { type: "exponential", delay: 60_000 },
23 },
24 })),
25 );
26 return { enqueued: created.length, jobIds: created.map((job) => job.id) };
27};
- The dedup key is the GitHub node id, so re-running the crawl doesn't queue the same repo twice.
attempts + exponential backoff is how rate-limited jobs come back later on their own (stage 4).
Stage 2: page through the stars
enqueue.ts fetches one GraphQL page (up to 100 stars), hands it to the queue, and moves on to the next cursor. It never holds the whole star list in memory, because the queue owns the payload from here on:
1const page = await client.getUserStarredReposMinimal({ login, first: pageSize, after });
2if (!page) return { data: "crawl-done", error: null };
3
4await enqueueStarredRepoEmbedJobs(
5 page.edges.map(({ node }) => ({
6 id: node.id,
7 owner: node.owner.login,
8 name: node.name,
9 description: node.description,
10 tags: node.tags,
11 languages: node.tags,
12 })),
13);
14
15const { hasNextPage, endCursor } = page.pageInfo;
16if (!hasNextPage || !endCursor) return { data: "crawl-done", error: null };
17if (pagesLeft !== undefined && pagesLeft <= 1) return { data: { nextCursor: endCursor }, error: null };
18
19return enqueueAllStarredRepos({ ...input, after: endCursor, pages: pagesLeft && pagesLeft - 1 });
Two knobs make it easy to try on a small slice first:
pages caps how many pages this run fetches. The result hands back nextCursor so a later run can resume from there.after starts from a saved cursor.
If GitHub rate-limits the list call itself, the same page is retried with exponential backoff (1 min, 2 min, 4 min, capped at 15) up to five times:
1} catch (caught) {
2 if (!isGithubRateLimited(caught)) throw caught;
3 if (retriesLeft <= 0) return { data: null, error: "429" };
4
5 const attempt = MAX_RATE_LIMIT_RETRIES - retriesLeft;
6 await sleep(Math.min(60_000 * 2 ** attempt, 15 * 60_000));
7 return enqueueAllStarredRepos({ ...input, retriesLeft: retriesLeft - 1 });
8}
For each repo, src/lib/embedding-gemmma/embed-repo.ts fetches the README, keeps the first 20 lines, and builds one short text document:
1export function buildRepoEmbedDocument(repo: EmbedRepoShape, readmeSummary: string | null): string {
2 const parts = [
3 `Repository: ${repo.name}`,
4 `Full name: ${repo.owner}/${repo.name}`,
5 repo.description ? `Description: ${repo.description}` : null,
6 repo.languages.length > 0 ? `Languages: ${repo.languages.join(", ")}` : null,
7 repo.tags.length > 0 ? `Tags: ${repo.tags.join(", ")}` : null,
8 readmeSummary ? `README:\n${readmeSummary}` : null,
9 ];
10 return parts.filter(Boolean).join("\n").slice(0, EMBED_TEXT_MAX_CHARS);
11}
Then it embeds that document with EmbeddingGemma in document mode (chapter 5 covers the model and why the mode matters):
1export async function embedRepo(repo: EmbedRepoShape): Promise<EmbedRepoResult> {
2 const client = createGitHubClient(await getGithubToken());
3 const readme = await client.getRepoReadme(repo.owner, repo.name);
4 const summary = clipReadmeSummary(readme?.content);
5 const text = buildRepoEmbedDocument(repo, summary);
6
7 await ensureOrtReady();
8 const { embedDocument, getEmbeddingModelId, getServerGemmaEmbedding } =
9 await import("@repo/gemma-embedding/node");
10 await getServerGemmaEmbedding({ dtype: readGemmaPrefs().dtype });
11 const vector = await embedDocument(text);
12
13 return { ...repo, summary, text, modelId: getEmbeddingModelId(), embedding: Array.from(vector) };
14}
One vector per repo, no chunking. Metadata plus the top of the README tells you what a repo is, which is what "find that starred repo about X" needs, and it fits comfortably in EmbeddingGemma's 2K-token context. The tradeoff is that something mentioned only deep in a README won't match. A chunk table is the natural next step if that ever matters.
A repo with no README still gets embedded from its name, description, and topics.
worker.ts pulls batches of 10 and handles each job on its own, so one bad repo doesn't fail the batch:
1export const starredRepoEmbedWorker = new Worker<StarredRepoEmbedJob>(
2 "repo-embed",
3 async (jobs) => {
4 const results: BatchResult[] = [];
5
6 for (let i = 0; i < jobs.length; i++) {
7 const { owner, name } = jobs[i]!.data;
8 patchEmbedActivity({ phase: "embedding", embed: { current: { owner, name } } });
9
10 try {
11 const embedded = await embedRepo(jobs[i]!.data);
12 const row = await upsertStarredEmbed(embedded);
13 const completed = getEmbedActivityStatus().embed.completed + 1;
14 patchEmbedActivity({ embed: { current: null, completed } }, row);
15 results.push({ status: "completed", value: { owner, name } });
16 } catch (caught) {
17 if (isGithubRateLimited(caught)) {
18 await pauseForRateLimit(asError(caught));
19 for (let j = i; j < jobs.length; j++) results.push({ status: "failed", error: asError(caught) });
20 return results;
21 }
22 results.push({ status: "failed", error: asError(caught) });
23 }
24 }
25 return results;
26 },
27 { store: starredRepoEmbedStore, batch: { size: 10 }, autoStart: false },
28);
The upsert is keyed on (owner, name), so re-embedding a repo replaces its row instead of duplicating it:
1await db
2 .insert(projectEnrichmentOutputs)
3 .values({ owner, name, type: "starred", description, summary, payload: { text }, modelId, embedding, embeddedAt: new Date() })
4 .onConflictDoUpdate({
5 target: [projectEnrichmentOutputs.owner, projectEnrichmentOutputs.name],
6 set: { description, summary, payload: { text }, modelId, embedding, embeddedAt: new Date() },
7 });
autoStart: false means the worker only runs when asked. POST /enrich/starred/run starts it, and there are worker/start, worker/pause, and worker/resume endpoints for manual control.
When a README fetch comes back rate-limited, the worker pauses itself for a minute and fails the rest of the batch without touching GitHub again:
1async function pauseForRateLimit(error: Error): Promise<void> {
2 starredRepoEmbedWorker.pause();
3 patchEmbedActivity({ phase: "waiting", list: { rateLimited: true }, message: "GitHub rate limited — pausing embed for 60s" });
4
5 await sleep(60_000);
6
7 starredRepoEmbedWorker.resume();
8 patchEmbedActivity({ phase: "embedding", list: { rateLimited: false }, message: "Resumed embed after rate-limit pause" });
9}
Those failed jobs aren't lost. Each has attempts: 5 with exponential backoff from 60 s, so Conveyor schedules them again later. Between the pause and the backoff, a large star list just finishes more slowly instead of hammering the API.
Workers run after the HTTP request that started them has ended, so there are no request headers to read a session from. The kickoff route stores the token for the process (src/lib/github-token.server.ts):
1let workerGithubToken: string | null = null;
2
3export function rememberGithubTokenForWorkers(token: string): void {
4 workerGithubToken = token.trim() || null;
5}
6
7export async function getGithubToken(): Promise<string> {
8
9
10
11}
On desktop, the UI passes the token from bindings.getGithubAccessToken() (chapter 2) in the kickoff body.
Every step calls patchEmbedActivity(...). That updates one status object (phase, counters, current repo, last error) and publishes it, together with the freshly upserted row if there is one, on the pub/sub bus. Chapter 4 streams that to the UI.
1export type EmbedActivityStatus = {
2 phase: "idle" | "listing" | "embedding" | "waiting" | "done" | "error";
3 login: string | null;
4 list: { after: string | null; fetchedTotal: number; enqueuedTotal: number; totalCount: number | null; rateLimited: boolean };
5 embed: { current: { owner: string; name: string } | null; completed: number; failed: number; lastError: string | null };
6 message: string | null;
7 updatedAt: string;
8};
"Done" is derived rather than tracked: once the list crawl has settled and the queue has nothing waiting or delayed, maybeMarkEmbedDone() flips the phase. It's called after each batch, when the worker emits drained, and on GET /activity, so a missed event can't leave the UI stuck on "embedding".
What survives a restart:
The status lives on globalThis for the same reason as the pub/sub bus: Vite HMR re-evaluates modules, and the worker and the SSE route must keep sharing one object.
- Embeddings run one at a time on a single shared model instance on the CPU. That's plenty for a personal star list, and it keeps memory predictable next to the UI.
- The first embed can trigger a model download if the bootstrap (chapter 5) hasn't run yet, so the first repo takes noticeably longer.
- README content past line 20 isn't indexed. That's deliberate for v1, as described in stage 3.