Embedding providers
ModusBrain ships with 16 embedding-provider recipes covering OpenAI, ZeroEntropy, Voyage, OpenRouter (single key, many hosted models), the major hosted alternatives, three local options, and a universal escape hatch (LiteLLM proxy). Runmodusbrain providers list to see the live registry; modusbrain providers explain --json emits a machine-readable matrix for agents.
This page is the human-readable counterpart: capability per provider, env-var setup, dimensions, cost, and known constraints.
Quick start
Init resolves your provider from env keys
As of v0.37,modusbrain init --pglite auto-detects which provider to use from your env vars. With OPENAI_API_KEY set, you get OpenAI. With ZEROENTROPY_API_KEY set, you get ZeroEntropy. If multiple provider keys are set, init fires an interactive picker. If no provider keys are set in a non-TTY context (CI, Docker build), init exits 1 with a paste-ready setup hint. Explicit flags (--embedding-model, --no-embedding) always win over env detection.
The resolved provider + dimensions get persisted to ~/.modusbrain/config.json atomically, so subsequent runs are deterministic across releases.
TL;DR table
Note on local providers. Ollama and llama-server have no required API key, so they don’t show up in env-detection auto-pick. Pick them explicitly with
--embedding-model ollama:<model> to avoid silently routing to a daemon that may not be running.
If first import fails
Ifmodusbrain import fails with expected N dimensions, not M, run modusbrain doctor. The output will print the exact modusbrain config set ... or modusbrain retrieval-upgrade command to repair the mismatch. You should not need to delete ~/.modusbrain. The bug-class that historically forced rm -rf recoveries is closed as of v0.37.
The doctor distinguishes two repair paths:
-
Empty brain (no embedded chunks yet) — drop and re-init at the right dim:
-
Non-empty brain — migrate cleanly with the supported reindex path:
Decision tree
- Cost-sensitive, English-only: Ollama (free, local) or Voyage (paid, best quality per dollar).
- Quality-first: Voyage
voyage-4-large(1024-2048 dims, ~3-4× more dense tokens than OpenAI tiktoken). - Code-heavy brain (gstack per-worktree, source repos): Voyage
voyage-code-3(1024 default; supports 256/512/1024/2048). Tuned on programming languages. Voyage publishes head-to-head numbers showing it outperforms their general flagships on code retrieval (voyageai.com/blog). For gstack’s per-worktree pglite-backed code brain, this is the right default — see Topology 3 indocs/architecture/topologies.md. - Reranking pair: ZeroEntropy
zerank-2is the hosted default intokenmaxmode (seedocs/ai-providers/zeroentropy.md). Voyagererank-2.5pairs cleanly with Voyage embeddings. - Local reranking (no API spend):
llama-server-rerankerrecipe (v0.40.6.1) — point modusbrain at your ownllama-server --rerankinginstance running Qwen3-Reranker or self-hosted ZeroEntropy weights. Samegateway.rerank()seam, $0 per call. Walkthrough indocs/ai-providers/llama-server-reranker.md. - One key for many hosted models: OpenRouter. Set
OPENROUTER_API_KEYand useopenrouter:<provider>/<model>for chat against GPT-5.2, Claude 4.x, Gemini 3, DeepSeek, and dozens more without juggling per-provider keys. Embedding catalog includes OpenAI, Google, Qwen, BGE-M3. - Enterprise compliance: Azure OpenAI (data residency + private endpoints) or self-hosted via llama-server / Ollama.
- China region: DashScope (Alibaba) or Zhipu (BigModel). DashScope’s international endpoint at
dashscope-intl.aliyuncs.com; overrideprovider_base_urls.dashscopefor the China endpoint. - OSS local, full control: llama-server (
llama.cpp) for any GGUF model; Ollama for the curated catalog. - Anything else: LiteLLM proxy. Run LiteLLM in front of any provider (Bedrock, Vertex, Cohere, Jina, Fireworks, etc.) and point modusbrain at it via
LITELLM_BASE_URL.
Per-provider details
OpenAI
Default. SetOPENAI_API_KEY. Models: text-embedding-3-large (3072 max, 1536 default), text-embedding-3-small (1536). Matryoshka via the dimensions field — modusbrain pins it from embedding_dimensions config so existing 1536-dim brains stay aligned across SDK upgrades.
Voyage AI
Best-in-class quality on the Voyage 4 family (Jan 2026 release). SetVOYAGE_API_KEY. Models: voyage-4-large, voyage-4, voyage-4-lite, voyage-4-nano, voyage-3.5, voyage-code-3 (code-tuned), voyage-finance-2, voyage-law-2, voyage-multimodal-3 (text + image).
Voyage 4 family shares an embedding space across all variants, so you can index with voyage-4-large and query with voyage-4-lite without reindexing. Dims: 256, 512, 1024, 2048. 2048 exceeds pgvector’s HNSW cap of 2000 — those brains fall back to exact vector scans (still correct, just slower).
For brains that index source code (gstack’s per-worktree pglite-backed code brain — see Topology 3 in docs/architecture/topologies.md), prefer voyage-code-3 over voyage-4-large. Voyage tunes it on programming languages and publishes head-to-head numbers vs their general flagships on code retrieval. Configure at install time:
modusbrain reinit-pglite --embedding-model voyage:voyage-code-3 --embedding-dimensions 1024 (PGLite) or follow docs/embedding-migrations.md (Postgres). modusbrain config set embedding_model is refused — the schema column has to resize.
modusbrain reindex --code will print a recommendation when run against a brain whose configured embedding model isn’t code-tuned; suppress with MODUSBRAIN_NO_CODE_MODEL_NUDGE=1 if you’ve intentionally chosen another model (single-vendor procurement, compliance, etc.).
Google Gemini
SetGOOGLE_GENERATIVE_AI_API_KEY (the AI Studio public API key). Model: gemini-embedding-001. Default 768 dims; Matryoshka up to 3072. Cheap.
For GCP service-account / Vertex AI auth (production deployments), see the v0.32.x follow-up — Vertex ADC is on the roadmap.
OpenRouter
Single OpenAI-compatible API for fan-out to OpenAI, Anthropic, Google, DeepSeek, Meta Llama, Qwen, and dozens of other hosted providers. One key, many models. SetOPENROUTER_API_KEY and use openrouter:<provider>/<model> (e.g. openrouter:openai/gpt-5.2, openrouter:anthropic/claude-sonnet-4.6).
Embedding: openai/text-embedding-3-small (1536d default, Matryoshka shrink to 512/768/1024). OR’s embedding catalog also includes text-embedding-3-large, google/gemini-embedding-2-preview, qwen/qwen3-embedding-8b, bge-m3 — opt in via --embedding-model openrouter:<id>. Pricing matches the upstream provider (OR adds a small markup).
Chat: every chat model OR proxies works through /v1/chat/completions. The recipe lists 8 curated entry points (GPT-5.2 family, Claude 4.5/4.6/4.7, Gemini 3 Flash Preview, DeepSeek); any other OR catalog ID also works. Tool-calling envelope is supported by the OR endpoint, but per-model capability varies — check https://openrouter.ai/models before counting on tools for a specific slug.
Optional env:
OPENROUTER_BASE_URL— point at a self-hosted OR-compatible proxy.OPENROUTER_REFERER(defaulthttps://modusbrain.ai) andOPENROUTER_TITLE(defaultmodusbrain) — attribution headers for OR’s leaderboard. Forks running modusbrain inside a different agent stack (OpenClaw deployments etc.) should set these so their traffic gets attributed to them, not modusbrain.
tool_use_id across crashes/replays). OR-routed Anthropic is rejected at submit time regardless of the recipe flag. If you want the price/availability story OR offers for tool-calling, use it for chat only and keep an Anthropic key for subagent work.
Azure OpenAI
Enterprise OpenAI behind Azure tenancy. Required env:AZURE_OPENAI_API_KEY, AZURE_OPENAI_ENDPOINT (e.g. https://my-resource.openai.azure.com), AZURE_OPENAI_DEPLOYMENT (the deployment name from your Azure portal). Optional: AZURE_OPENAI_API_VERSION (defaults to 2024-10-21).
Unlike vanilla OpenAI, Azure uses api-key: header (not Authorization: Bearer) and a templated URL with ?api-version= query param — modusbrain handles both via the recipe’s resolveAuth + resolveOpenAICompatConfig overrides.
Models: text-embedding-3-large, text-embedding-3-small, text-embedding-ada-002 (your Azure deployment must serve the requested model).
MiniMax (海螺AI)
SetMINIMAX_API_KEY. Optional MINIMAX_GROUP_ID for org-scoped accounts. Model: embo-01 (1536 dims).
MiniMax’s API takes a type: 'db' | 'query' field for asymmetric retrieval. v0.32 routes everything as type='db' (symmetric retrieval — same vector space for indexing and queries). Asymmetric query support is a v0.32.x follow-up.
DashScope (Alibaba)
SetDASHSCOPE_API_KEY. International endpoint at dashscope-intl.aliyuncs.com by default; override provider_base_urls.dashscope for the China endpoint. Models: text-embedding-v3 (current; Matryoshka 64-1024 dims), text-embedding-v2.
CJK-dominant content tokenizes denser than OpenAI tiktoken; modusbrain declares chars_per_token: 2 so the batch pre-split leaves headroom.
Zhipu AI (BigModel)
SetZHIPUAI_API_KEY. Models: embedding-3 (current; Matryoshka 256-2048 dims), embedding-2. v0.32 default is 1024 (HNSW-compatible). The 2048-dim option works but falls into the exact-scan branch (see Voyage 4 Large note above).
Ollama (local)
No env required — Ollama runs unauthenticated locally. OptionalOLLAMA_BASE_URL (default http://localhost:11434/v1) and OLLAMA_API_KEY (for auth-enabled deployments).
Recipe ships with nomic-embed-text (768d, recommended), mxbai-embed-large (1024d), all-minilm (384d). modusbrain providers test --model ollama:nomic-embed-text smoke-tests the local install.
llama-server (local, llama.cpp)
llama.cpp’s llama-server --embeddings endpoint. No env required. Optional LLAMA_SERVER_BASE_URL (default http://localhost:8080/v1) and LLAMA_SERVER_API_KEY.
User-driven models: launch llama-server with --model <gguf-path> --embeddings, then run modusbrain init --embedding-model llama-server:<your-id> --embedding-dimensions <N>. The recipe refuses the implicit shorthand --model llama-server because there’s no canonical first model.
LiteLLM proxy (universal escape hatch)
Run LiteLLM in front of any provider — Bedrock, Vertex, Cohere, Jina, Fireworks, OctoAI, etc. The proxy normalizes everything to the OpenAI-compatible API; modusbrain points at the proxy viaLITELLM_BASE_URL and proxies the call.
This is the catch-all for “my provider isn’t in the list above.” Set up LiteLLM, then modusbrain init --embedding-model litellm:<your-model-id> --embedding-dimensions <N>.
Choosing dimensions
Three numbers matter:- Provider’s native dims: each model has a “true” output dim (e.g. OpenAI
text-embedding-3-largeis 3072 native). - Matryoshka reductions: most modern providers let you request a smaller vector via the
dimensionsfield. - HNSW cap: pgvector’s HNSW index supports up to 2000 dims. Brains above that fall back to exact vector scans (slower but correct; modusbrain handles the SQL automatically via
chunkEmbeddingIndexSqlinsrc/core/vector-index.ts).
My provider isn’t listed
Four options:- Use OpenRouter when the provider/model is available through OR’s OpenAI-compatible API (covers most hosted chat models + a growing embedding catalog).
- Use LiteLLM proxy (above) — the universal escape hatch. Works for 100+ providers.
- Open a feature request at github.com/thebuildceo/modusbrain/issues with the provider’s API docs URL and a setup snippet. Recipes are ~30-40 lines of TypeScript.
- Submit a recipe: clone, copy
src/core/ai/recipes/voyage.tsas the gold-standard openai-compat template, register insrc/core/ai/recipes/index.ts, add a per-recipe smoke test undertest/ai/recipe-<name>.test.ts. The recipe contract test (test/ai/recipes-contract.test.ts) and IRON RULE regression test pin the structural invariants.
Switching providers on an existing brain
Embedding dimensions are baked into the schema atmodusbrain init time. As of v0.37.11.0, modusbrain config set embedding_model and modusbrain config set embedding_dimensions are refused — the schema column has to resize alongside the config, and config set only touches the config row.
The supported paths:
- PGLite (default install):
modusbrain reinit-pglite --embedding-model <provider>:<model> --embedding-dimensions <N>— one-command wipe-and-reinit that preserves every other config field (chat model, expansion model, API keys), backs up the prior brain to<path>.bak, runsmodusbrain initwith the new flags, and re-syncs your brain repo. Add--no-syncto skip the resync,--yesto skip the TTY confirmation,--jsonfor scripts. - Postgres (Supabase / self-hosted): follow the SQL recipe in
docs/embedding-migrations.md(drop the HNSW index, ALTER COLUMN TYPE, clear stale embeddings, recreate the index conditionally, thenmodusbrain init --supabase --embedding-model X --embedding-dimensions Nto update the file plane and re-embed).
modusbrain doctor 8c “alternative_providers” surfaces unconfigured providers whose env is already set — useful when you’ve configured OpenAI but also have e.g. VOYAGE_API_KEY exported and want to know you can switch without extra setup.
modusbrain doctor 8c “alternative_providers” surfaces unconfigured providers whose env is already set — useful when you’ve configured OpenAI but also have e.g. VOYAGE_API_KEY exported and want to know you can switch without extra setup.