# Search Index Sync Engine

> Real-time & bulk search indexing

Keeping search indices in sync with a live database — in bulk and in real-time.

- **Page:** https://siddharthdeshpande.com/projects/search-sync
- **Built at:** Apra Labs
- **Tech stack:** C# / .NET 8, Cognitive Search, Service Bus, Durable Functions

## Impact

- **Real-time** — incremental sync
- **Bulk** — full re-indexing
- **Dual** — sync modes
- **Zero** — downtime deploys

## Problem

- Search results always slightly stale
- Full re-indexing painfully slow and manual
- Record updates silently not reaching search
- Users unable to find recently added records
- Schema changes require extended downtime
- No retry mechanism for failed sync events

## Solution

- Dual-mode sync engine built from scratch
- Real-time incremental sync via Service Bus
- Bulk orchestration via Azure Durable Functions
- Index versioning enabling zero-downtime deploys
- Dead letter queue for failed event recovery
- Automatic schema migration with version control

## Outcome

- Real-time search sync fully operational
- Zero-downtime schema migrations in production
- No silently dropped changes ever again
- Dual sync modes running for all indexes
- Failed events automatically retried and recovered
- Search freshness under 5 seconds end-to-end

## How it works

1. **Real-time Sync** — Service Bus captures database change events. Each event triggers incremental index updates within seconds of the source change.
2. **Bulk Sync** — Durable Functions orchestrate full re-indexing with checkpoint and retry logic. Field mapping transformations handle schema differences between source and search index.
3. **Operations** — Sync lag and failure monitoring ensures no changes are silently dropped. Index versioning handles schema evolution without downtime.

## Key decisions

- **Dual-mode sync, not one or the other** — Bulk mode handles full re-indexing and schema migrations. Real-time mode handles individual record changes within seconds. Both are necessary — one without the other always leaves gaps.
- **Index versioning for zero-downtime schema changes** — Schema updates create a new index version, populate it in the background, then swap aliases atomically. Users never hit a partially-updated or stale index.

## What I'd change

- **Dead letter handling needs automation** — Failed sync events go to a dead letter queue, but recovery is manual. An automated retry-and-reconcile process would reduce operational overhead.
- **Per-entity sync strategies earlier** — Initially used one strategy for all entity types. Making this configurable per entity from the start would have simplified later tuning.
