Skip to content

Python CLI + library · local SQLite · offline by default

Index only what changed.Reuse everything else.

Steadlith maintains local RAG indexes when source documents change. It gives chunks stable identities, shows a dry-run plan, reuses cached embeddings, then publishes one transactional SQLite update.

Plan first. No network or index write unless you explicitly allow it.

Illustrative output from the Steadlith plan command
SCLI dry run
NO WRITES
Index plancomplete desired corpus
OperationChunksWhat happens
add2embed cache misses
keep187reuse existing identities
move1reuse content at a new position
delete1requires explicit approval
$steadlith plan
2 embeddings neededno writesno network
01 · SourcesMarkdown and text files
02 · Chunkstable content identities
03 · Planadd, keep, move, delete
04 · Embedreuse cached vectors
05 · Publishone SQLite transaction

The cost of unstable boundaries

A small edit should not force a full re-index.

Offset-based chunking can shift every downstream boundary after an early insertion. The text may be unchanged while its chunk hashes and embeddings are not.

Offset chunks
Where does this fixed window begin now?
Steadlith
Which content identities actually changed?

Steadlith reports a concrete manifest delta. It does not promise that every edit changes only a fixed number of chunks.

One explicit state model

Every transition stays inspectable.

Chunk identity, cache identity, manifest state, and index publication remain separate so the plan can explain what will happen before it happens.

StageMechanismObservable behaviorDefault
IdentityContent-defined chunksNormalize words, place Rabin boundaries, and hash canonical content with versioned parameters.Deterministic
PlanManifest diffClassify add, keep, move, and delete operations before any provider call or state write.Read only
CacheContent-addressed embeddingsReuse vectors by chunk, provider, model, and embedding parameters.Local SQLite
ApplyTransactional publicationReject stale plans, tombstone removals, and publish one validated index generation.Explicit approval
VerifyManifest and index checksCompare generations, record digests, active rows, and Merkle roots.Offline
QueryPluggable embeddingsUse deterministic lexical retrieval or an explicitly selected learned provider.No credentials
Read the complete documentation

One workflow, three explicit stages

Plan before writing. Verify after.

The CLI exposes the same planner and index service available from the Python API.

Plan

Inspect the complete delta

Resolve the desired corpus and price only known cache misses without changing state.

terminal
steadlith plan
Apply

Publish one generation

Reuse cached vectors, embed misses, and approve destructive operations explicitly.

terminal
steadlith index --allow-delete
Verify

Check committed state

Confirm that the active index, manifest, generation, and record digests agree.

terminal
steadlith verify

Small, explicit trust boundary

Pure identities. Effects at the edge.

Chunking and planning stay deterministic. Files, credentials, providers, and SQLite enter through explicit application boundaries.

Read the architecture
  1. 01

    Chunk

    Normalize words and place content-defined Rabin boundaries.

  2. 02

    Identify

    Hash canonical content with the versioned chunking recipe.

  3. 03

    Plan

    Compare manifests and resolve reusable embedding identities.

  4. 04

    Publish

    Apply one transaction, tombstone removals, and verify state.

No network by defaultNo implicit deletionsNo silent migrationsNo remote backend claim

Install the release

Evaluate Steadlith on a real corpus.

Install v1.0.0, start with the offline provider, inspect the plan, and measure churn before connecting a paid embedding service.

Open the quick start
terminal
python -m pip install steadlith
steadlith init
steadlith plan

Apache-2.0, Python 3.10+, local SQLite reference backend

View on PyPI