v0.4 · demo dataset

An enterprise context pipeline that makes LLM answers defensible.

Context Synthesizer unifies Slack, Jira, Google Drive, and Notion into a single semantic layer — with hybrid retrieval, parent-child chunking, entity graphs, and continuous evaluation of faithfulness, recall, and groundedness.

documents indexed
244,606
unified schema
6 systems
median freshness
94 s
retrieval P@8
0.912

Ingestion surface

Four systems, one canonical schema

Slack

● healthy
docs
148,213
last sync
42s ago
throughput
1.2k msg/min
scope
23 channels · 4 workspaces

Jira

● healthy
docs
24,817
last sync
1m ago
throughput
180 issues/min
scope
9 projects · 6 boards

Google Drive

● syncing
docs
62,194
last sync
3m ago
throughput
420 files/min
scope
17 shared drives

Notion

● degraded
docs
9,382
last sync
12m ago
throughput
60 pages/min
scope
5 teamspaces

Positioning

Why naive RAG fails inside a real company

Enterprise questions cross tools, span time, and depend on who is asking. Retrieval that treats every chunk as an island cannot answer them.

Naive RAG

What breaks

  • Chunks lose their document

    Retrieved slices arrive with no parent, no author, no permissions — the model hallucinates ownership.

  • Single-source retrieval

    Vector-only search misses exact identifiers (`ATLAS-874`, `/v2/search`) that BM25 nails.

  • No entity model

    Systems don't know that #mobile-auth in Slack, ATLAS-874 in Jira, and RFC-024 in Notion describe the same thing.

  • No eval loop

    Nobody knows the retrieval regressed — until a support ticket.

Context Synthesizer

How this system answers

  • Parent-child linkage

    Every chunk carries its parent summary and metadata; citations resolve back to the exact section.

  • Hybrid + rerank

    BM25 + dense retrieval fused with RRF, then a cross-encoder rerank on the top-40.

  • Semantic graph

    Entities are extracted once and reused; cross-system context reconstructs on retrieval.

  • Continuous evaluation

    Ragas + Phoenix score every trace: faithfulness, recall, groundedness, precision.

How it works

The pipeline, end to end

  1. 01

    Ingest

    Connector workers pull deltas from Slack, Jira, Drive, Notion.

  2. 02

    Normalize

    Map to canonical envelope; preserve ACLs, authors, timestamps.

  3. 03

    Chunk

    Heading-aware split with parent-child linkage.

  4. 04

    Embed

    bge-large for dense; BM25 index for lexical.

  5. 05

    Retrieve

    RRF fusion, top-40 → cross-encoder rerank → top-8.

  6. 06

    Graph

    Entity extraction stitches cross-system context.

  7. 07

    Synthesize

    Grounded answer with numbered source citations.

  8. 08

    Evaluate

    Ragas scores every trace; Phoenix stores the run.

Live dashboard · demo data

Retrieval, evaluation, and ingestion health

IngestionRetrievalRerankerGraphEval loopNotion sync

Retrieval precision

0.912+1.4%

Faithfulness

0.947+0.6%

Context recall

0.881-0.3%

Answer groundedness

0.934+0.9%

Ingestion freshness

94s-12s

Evaluation trends · 14 days

Precision · faithfulness · recall

precisionfaithfulnessrecall

Source contribution

Citations by system · 7 days

Ingestion health

By source

SourceDocsFreshnessError rateLast syncStatus
slack148,21398%0.20%42s● healthy
jira24,81796%0.40%1m● healthy
drive62,19491%1.10%3m● syncing
notion9,38278%3.20%12m● degraded

Retrieval pipeline

Query → top-8

query1BM25412vector512RRF40rerank8

latencies: p50 620ms · p95 1.4s

Recent traces

Last 5 runs

phoenix://traces
  • trc_9f2aok

    What changed in Project Atlas over the last 3 sprints?

    bm25 → vector → rrf → rerank → graph → synthesize

    842ms
    f0.96 · 7 cites
  • trc_9f19ok

    Owner of the payments idempotency-key spec?

    bm25 → vector → rrf → rerank → synthesize

    611ms
    f0.98 · 3 cites
  • trc_9f0cok

    Open blockers on mobile auth this week

    bm25 → vector → rrf → rerank → graph → synthesize

    1204ms
    f0.93 · 9 cites
  • trc_9ef7low_faithfulness

    Q3 SLA breach postmortem — root cause

    bm25 → vector → rrf → rerank

    1583ms
    f0.71 · 2 cites
  • trc_9ee1ok

    Latest rate limit values for /v2/search

    bm25 → vector → rrf → synthesize

    498ms
    f0.99 · 4 cites

Failed retrievals

Needs attention

  • fq_412permission_filtered

    What did legal decide about the EU data residency clause?

    3 candidate chunks excluded by ACL (legal-internal)

  • fq_411stale_index

    Design review notes for onboarding v4

    Notion partition last synced 42m ago (SLA: 5m)

  • fq_408no_high_confidence_source

    Which vendor was picked for the observability RFP?

    Top rerank score 0.41 (threshold 0.55)

Semantic entity graph

Cross-system topology

42 nodes · 118 edges
Project Atlasmobile-authPaymentsRate limitsQ3 SLAEU residency

Top entities · 24h

By retrieval weight

  • Project Atlas

    project

    42
  • mobile-auth

    workstream

    31
  • API rate limits

    topic

    27
  • Payments

    team

    24
  • Q3 SLA breach

    incident

    19
  • EU data residency

    policy

    17

Ingestion freshness

Median seconds behind source

SLA · 120s

Interactive demo

Ask across every system

Pick a real enterprise question. See what was retrieved, from where, and how the synthesized answer scored.

↵ run

Retrieved context · 4 chunks

  • jira[1] ATLAS-812 · Partition ingestion by source
    rerank 0.91

    Split the monolithic worker into per-source queues. Backpressure isolated; Notion no longer starves Slack.

    retrieval
    0.82
    trust
    0.94
    freshness
    0.97
    · closed · sprint 41 · @m.ito
  • notion[2] RFC-019 · Parent-child chunking
    rerank 0.88

    Long Drive docs are chunked at heading boundaries; parent doc summary is co-retrieved for context reconstruction.

    retrieval
    0.78
    trust
    0.92
    freshness
    0.90
    · approved · sprint 42 · @s.chen
  • drive[3] retrieval-eval-sprint43.pdf
    rerank 0.93

    RRF(k=60) over BM25 + bge-large; cross-encoder rerank on top-40 → top-8. Precision@8 = 0.912.

    retrieval
    0.86
    trust
    0.90
    freshness
    0.88
    · shared · sprint 43 · platform-search
  • slack[4] #atlas-standup · mobile-auth descope
    rerank 0.85

    Given the Q3 SLA postmortem, moving mobile-auth to sprint 44. Payments team owns the auth-refresh spike.

    retrieval
    0.75
    trust
    0.88
    freshness
    0.99
    · thread · 6 replies · @j.kim

Synthesized answer

Across sprints 41–43, Project Atlas shifted from a monolithic ingestion worker to a partitioned queue-per-source model [1], introduced parent-child chunking for long Drive documents [2], and adopted RRF fusion with a cross-encoder reranker [3]. The mobile-auth workstream was descoped to sprint 44 after the Q3 SLA postmortem [4]. Retrieval precision improved from 0.87 to 0.91 as a result.

System architecture

Eight layers, one contract

Each layer has a narrow contract with the next. Traces flow forward; evaluation flows backward.

  1. L1

    01/8

    Ingestion connectors

    Per-source workers with delta pagination, backoff, and dead-letter queues. Slack, Jira, Drive, Notion.

    Slack APIJira APIDrive APINotion API
  2. L2

    02/8

    Normalization & metadata

    Every record maps to a 15-field canonical envelope: source_type, ids, authors, teams, permission_tags, timestamps, trust_score.

    PydanticOpenTelemetry
  3. L3

    03/8

    Parent-child chunking

    Heading-aware splits keep chunks tight; parent doc summary is co-indexed so retrieved slices arrive with their context.

    semchunktiktoken
  4. L4

    04/8

    Hybrid retrieval

    BM25 (Postgres FTS) and dense vectors (bge-large in pgvector + Qdrant) fused with reciprocal rank fusion.

    Postgres · pgvectorQdrantRRF
  5. L5

    05/8

    Cross-encoder reranking

    Top-40 candidates reranked by ms-marco cross-encoder to top-8 with score-gated fallback to no-answer.

    cross-encoderscore gating
  6. L6

    06/8

    Semantic entity graph

    Entities extracted at ingest and linked across systems. Retrieval walks the graph to reconstruct cross-tool context.

    spaCyNeo4j-compatible
  7. L7

    07/8

    Evaluation & telemetry

    Ragas scores faithfulness, recall, and precision on every trace. Phoenix stores runs for drift analysis.

    RagasArize Phoenix
  8. L8

    08/8

    Secure source attribution

    Permission tags travel with every chunk. Answers cite specific documents; ACL-filtered candidates are logged.

    ACL propagationaudit log

Why this project

What this demonstrates to an AI engineering hiring manager

Context Synthesizer is deliberately shaped like real enterprise infrastructure: a messy multi-source ingestion problem, a retrieval stack that combines lexical and semantic signals, an evaluation loop that catches regressions, and a UI that treats citations and permissions as first-class.

Advanced RAG

Hybrid retrieval, RRF fusion, cross-encoder reranking, parent-child chunking, score-gated no-answer.

Data engineering

Multi-source ingestion, delta pagination, ACL preservation, canonical schemas, dead-letter handling.

Evaluation

Ragas metrics wired into every trace: faithfulness, recall, precision, groundedness.

Observability

Phoenix traces, per-stage latency, drift dashboards, failure taxonomy.

Semantic modeling

Entity extraction, cross-system linkage, graph-augmented retrieval.

AI product engineering

Recruiter-facing dashboards, source attribution, permission-aware answers.

Stack

  • FastAPI
  • Postgres · pgvector
  • Qdrant
  • Ragas
  • Arize Phoenix
  • Slack API
  • Jira API
  • Google Drive API
  • Notion API
  • Hybrid Search
  • Knowledge Graph