Batch RAG for high-volume ingest and bulk query scenarios. Covers OpenAI
Batch API (50% discount, 24h SLA), Anthropic Message Batches API, Voyage
AI and Cohere batch embeddings, ingestion-time vs query-time batching,
async/Ray parallelism, parallel writes to vector DBs, and rate-limit
coordination across workers.
USE WHEN: user mentions "batch API", "OpenAI batch", "Anthropic batches",
"bulk embedding", "Ray embeddings", "parallel ingest", "batch RAG"
DO NOT USE FOR: streaming single-query latency - use `rag-production`;
cost-dashboards - use `cost-allocation`;
LLM gateway routing - use `llm-gateway`