Hybrid sparse-dense retrieval, now generally available

The reference layer your RAG pipeline needs

Normalize, chunk, embed, and retrieve across heterogeneous corpora. One API, compatible with LangChain, LlamaIndex, and any custom stack.

ragref retrieve
$ curl ragref.com/v1/retrieve \ -d '{"corpus":"support-docs", "query":"reset 2FA device","rerank":true}' # 200 OK · 41ms { "matches": [ { "score": 0.914, "doc": "security/2fa.md" }, { "score": 0.877, "doc": "account/devices.md" } ], "latency_ms": 41 }
Drops into the stack you already run
10Btokens per corpus
100kqueries per second
<50msp99 retrieval latency
99.95%indexing uptime

Built for production RAG

Retrieval quality and index durability, without operating your own vector infrastructure.

Retrieval

Hybrid search that holds up out of domain

BM25 sparse and dense vectors are fused with reciprocal rank fusion, then optionally passed through a cross-encoder reranker before anything reaches your context window.

  • Sparse + dense fusion with tunable weighting
  • Cross-encoder reranking on retrieved candidates
  • Citation provenance on every match
dense0.72
sparse0.58
fused0.91
reranked0.96
Chunking

Semantic chunking that preserves context

Paragraph, sentence, and token-boundary strategies with semantic overlap detection, so a split never severs a reference from the claim it supports.

  • Configurable boundaries and overlap
  • Structure-aware splitting for code and tables
  • Deterministic, replayable ingestion
Two-factor authentication adds a second step at sign-in. To reset a lost device, an account owner opens Settings, selects Security, and confirms identity by email before a new authenticator can be enrolled.
overlappreserved

A small, predictable API

Everything is scoped to a corpus. Everything returns provenance.

Corpora
POST/v1/corporaCreate a new document corpus
POST/v1/corpora/{id}/ingestIngest and chunk documents
GET/v1/corpora/{id}/statsCorpus metadata and index health
Retrieval
GET/v1/retrieveHybrid sparse-dense retrieval
POST/v1/rerankCross-encoder reranking
GET/v1/embedEmbeddings, OpenAI-compatible

Live in five lines

Install the client, create a corpus, ingest a directory, and retrieve. No infrastructure to stand up.

# pip install ragref from ragref import RAGRefClient client = RAGRefClient(api_key="rr_live_...") corpus = client.corpora.create(name="support-docs", chunking="semantic") corpus.ingest(path="./docs/", recursive=True) results = client.retrieve( corpus_id=corpus.id, query="How do I reset my 2FA device?", top_k=8, rerank=True, )

Pay for what you retrieve

Indexing is free at every tier. Billing tracks retrievals and stored tokens, nothing else.

PRO · MOST POPULAR
$149/mo
100M tokens indexed
  • 500,000 retrievals/mo
  • Unlimited corpora
  • Hybrid search + reranking
  • Webhook events
  • Priority email support
  • Retrieval analytics
Get API key
STARTER
$0/mo
Up to 1M tokens indexed
  • 10,000 retrievals/mo
  • 2 corpora
  • Basic chunking
  • Community support
Start free
ENTERPRISE
Custom
Unlimited scale
  • Unlimited tokens & retrievals
  • Private VPC deployment
  • Custom embeddings
  • SLA + dedicated CSM
Talk to us

Ship retrieval you can trust

Spin up a corpus and run your first hybrid query in minutes.

Get API key