Lante
Company Brain for logistics ops — WhatsApp exports to a cited GraphRAG audit trail
- Role
- Sole developer
- Stack
- PythonFastAPINext.jsPWATypeScriptPostgresQdrantFalkorDBDocker
Lante has no public always-on URL — the stack is spun up on request. Get in touch, I'll approve access, we'll schedule a time, and I'll turn the project on for a live walkthrough. Request a walkthrough
Problem
In many work settings, keeping multiple projects straight at once is hard — and it gets harder as each project grows in moving parts, stakeholders, and exceptions.
Most people do not live in a project-management tool. They live in email and WhatsApp, especially in LATAM. Updates, handoffs, and “where did we leave this?” moments scatter across threads instead of sitting in one place you can trust.
Lante ingests what people already send and receive on those channels, then maintains a current-state view per project — not just search, but “what is the status of this shipment / order / client right now?”
We focused the first domain on freight logistics on purpose: multiple orders in flight, protocol-heavy workflows, and sudden changes (delays, re-routes, customs holds) are normal daily work. That sector stress-tests whether the system can stay accurate when volume and volatility are both high.
Solution
Upload a WhatsApp export (or a PDF). Once it is ingested — and confirmed in Bandeja if the routing was ambiguous — you can ask, in Spanish, ¿Quién prometió la entrega del pedido #8890? and get citations pointing at specific messages, plus a timeline of the Friday promise, the slip, and the Monday re-promise. The seeded #8890 chat is one of the ambiguous ones: it spans several topics, so it does go through review.
The web PWA is the shipping client. A shared TypeScript package (apps/shared) is the API client for web (and the Expo app, which typechecks only). FastAPI serves the API. Docker Compose runs Postgres, Redis, Qdrant, FalkorDB, MinIO, the API, a Postgres-polling ingest worker, and a rules worker.
Walkthrough and private repo access are available on request — there is no public always-on URL. See How to see it.
Tech stack
| Layer | What actually ships |
|---|---|
| Language / API | Python 3.11, FastAPI, Pydantic, SQLAlchemy, Alembic (12 revisions) |
| Web | Next.js 15, React 18, TypeScript; PWA via manifest.json, sw.js, install prompt |
| Shared client | apps/shared — typed fetch wrappers used by the PWA |
| Datastores | Postgres 16 (system of record), Redis 7 (rate limits and short-lived tokens — not the ingest queue), Qdrant (chunk vectors), FalkorDB (entity graph over the Redis protocol), MinIO (raw uploads) |
| LLM | Default gpt-4o-mini at extraction temperature 0; embeddings text-embedding-3-small (1536-d); optional per-workspace Qwen via DashScope |
| Eval / NLP extras | Hand-annotated JSONL gold sets, rapidfuzz, pdfplumber for PDFs |
| Mobile / extras | Expo 52 app typechecks in CI; Telegram bot module exists; telegram: false on every plan in packages/billing/tiers.py |
Redis is in the stack. The worker does not consume a Redis queue — it polls the Postgres Job table.
Skills this demonstrates
| Skill | Evidence in the repo |
|---|---|
| Applied NLP / extraction | Spanish LLM prompt + JSON schema, heuristic _mock_extract fallback, commitment regex overlay, rapidfuzz resolver (threshold 85) |
| Information retrieval | Hybrid retriever: Qdrant vector search + FalkorDB Cypher for commitment edges, lexical +0.1 fusion, citations with sender and timestamp |
| Evaluation design | 93-row extraction gold/test split, 48 negatives, 20 phenomena (≥3 each), circular-fixture hygiene test, retrieval gold set of 21 questions with sender/date citation scoring |
| LLM product engineering | Tool-using ops copilot, deterministic synthesis for delivery-promise questions, bilingual input blocklist, Spanish refusals, optional NVIDIA NeMo topic/content guards (off by default) |
| Multi-tenant backend | require_workspace_access / WorkspaceDep; test_route_auth_coverage.py enumerates live routes; per-workspace payload check on every vector hit |
| Full-stack product | Inbox HITL, timelines, copilot UI with citation cards, audit CSV/PDF export, PWA |
| Production judgement | Rate limits fail open; embedding-dimension mismatch raises; demo mode refuses to boot without closed registration + read-only email; CI blanks LLM keys on the unit lane |
This does not demonstrate statistical inference, forecasting, or training a model from scratch (that’s Sol).
Architecture
Two indexes on purpose. Vector search answers “what was said about freight.” Graph traversal answers “who promised what.” Retrieval queries both.
| Layer | Choice | Why (and where) |
|---|---|---|
| Chunking | 800 characters, 100 overlap | packages/ingestion/chunker.py |
| Extraction | gpt-4o-mini, temperature 0, extract_entities_es.md | Informal Colombian Spanish; heuristic path if the LLM is empty |
| Embeddings | text-embedding-3-small (1536-d); mock 384-d without a key | packages/indexing/embedder.py — collection dim is fixed at creation |
| Vector store | Qdrant | One shared collection. document_id uses a server-side filter; workspace_id is checked on each returned payload (see below) |
| Graph | FalkorDB | Cypher; smaller ops footprint than Neo4j |
| Entity resolution | rapidfuzz token_sort_ratio, threshold 85 | “Fernando” / “Fernando Quintero” / “don Fernando” |
| Serving | FastAPI + Postgres-polling worker | No message broker; ingest is minutes-scale |
Where tenant isolation actually happens on vector search. All workspaces share one Qdrant collection. document_id is pushed down as a server-side Filter, but workspace_id is enforced in application code — vector_writer.search() drops any returned point whose payload workspace_id does not match. Isolation holds, but two things follow honestly: another tenant's nearest neighbours can consume part of the top_k budget before being discarded, and the guarantee lives in one Python line rather than in the database. A payload index with a server-side must clause is the correct fix and is not done.
Human-in-the-loop is conditional, not universal. should_queue_for_review() sends a document to Bandeja when it splits into multiple segments, when more than one topic is detected, when a routing suggestion scores below 0.7 confidence, or when the workspace sets require_timeline_approval (default false). A single-topic, high-confidence upload routes straight to a timeline. The claim is that the ambiguous cases are queued — not that a human approves everything.
Measurement
Two harnesses are worth quoting. A third file, agent_eval.jsonl, has three cases and still mentions the old fixture’s Carlos / #4521 — it is not a published metric.
Extraction (offline)
Hand-annotated 93 rows (62 gold / 31 test) over 7 co_logistics_*.txt WhatsApp exports. Twenty phenomena, each tagged at least three times. 48 rows have empty entities and relations (deliberate negatives). Every row is "synthetic": true.
Required entities: 43 on gold, 63 across both splits. Relations: 8 total (4 per split).
Reproduced this session, no API key:
python scripts/run_extraction_eval.py --split gold --predictor heuristic
| Predictor | Entity F1 (micro) | Relation F1 (micro) |
|---|---|---|
| Heuristic regex + mock extract (gold, 62 rows) | 0.692 (P 0.771 / R 0.628; 27 TP, 8 FP, 16 FN) | 0.143 (1 TP, 9 FP, 3 FN) |
gpt-4o-mini, temperature 0 (gold, recorded in datasets/README.md) | 0.933 (0.923–0.944 over 3 runs) | 0.800 (same across those 3 gold runs) |
The 9 false relations vs 1 true one is why the LLM path exists. On the same heuristic run, colombian_logistics_term entity F1 is 0.286 and negative_ack is 1.000; type confusion is Location→Person 3, Organization→Person 2.
LLM entity/relation numbers above are not re-run in this pass (they need OPENAI_API_KEY). They are the figures recorded with the corpus in packages/evaluation/datasets/README.md.
Retrieval (live stack)
retrieval_gold_set.jsonl — 21 questions covering all seven conversations. Scoring: fraction of expected_contains substrings in the answer, and whether a citation matches expected_citation_sender and the same calendar day as expected_citation_timestamp. Gate in scripts/run_eval.py: retrieval ≥ 0.80, citation ≥ 0.70. Needs Docker + a seeded workspace + an API key. No number from that gate is quoted here.
Honesty about the numbers
Relation F1 is under-powered. Eight relations total. On the test split, two LLM runs of the same prompt scored 0.571 and 1.000. Do not treat gold 0.800 as a stable capability number.
Single annotator, synthetic corpus. No inter-annotator agreement. Treat F1 as an upper bound. Real anonymized exports would set "synthetic": false in the same schema.
The old retrieval gold set was circular. packages/evaluation/gold_set.jsonl still exists to reproduce old numbers; it is not the default gate.
Failure modes
Observed on this corpus / this code:
- Hedges. “si acaso llega el viernes, todavia no es seguro por el clima” is tagged
hedged_noncommitment(3 rows).commitment_extractor.pystill fires onentrega-adjacent regex. Inventing a promise is worse than missing one. - Superseded promises are appended.
#8890Friday then Monday: both edges can exist.valid_fromis set on relations (temporal_edges.py); nothing marks the first as void. - Accents. Gold names are verbatim (
trancón,Julián). Matching normalizes accents in eval; a “fix” to gold strings would still be wrong. - Heuristic type confusion. Capitalized Spanish tokens default to Person in
_mock_extract— hence Location→Person / Organization→Person on the baseline. - Implicit subjects. “hay un trancon verraco en la via al Llano” is annotated with a location and no relation, because the extractor sees
msg.body, notsender: body. - The deterministic answer is tuned to this corpus.
answer_synthesis.pyfires on any delivery-promise question naming an order of three or more digits, but the two-date sentence (“first Friday, then rescheduled to Monday”) is keyed on the literal wordsviernesandlunes. It is a correct, cheap answer for the seeded chats and would need real date parsing to generalise.
Production concerns
| Concern | Implementation |
|---|---|
| Multi-tenant isolation | Path/body workspace_id routes must call the membership helper. tests/unit/test_route_auth_coverage.py walks the live app. Deny-behavior is in test_tenant_isolation.py (requires_infra, not the offline 431) |
| Rate limiting | Redis fixed-window: login 5/min/IP, register 3/hour/IP, LLM/copilot 30/min per identity (/v1/consulting/chat included). Fails open if Redis is down |
| Prompt injection | Two layers on the copilot path: a 500-character cap plus a bilingual blocklist (security_blocklist.py — Spanish and English instruction-override, system-prompt extraction, code fences, SQL-injection asks), then an optional NeMo NIM topic check (nemo_guardrails_enabled defaults false) and an output-safety pass. A separate four-phrase English check sits in guardrails.py. Refusals are Spanish |
| Cost | Per-workspace OpenAI vs Qwen. Copilot 30/min. Delivery-promise questions can skip the LLM via answer_synthesis.py |
| Migrations | Alembic, 12 revisions; migrate container must exit 0 before api starts |
| Embedding dim | Mock 384-d vs OpenAI 1536-d; embedding_dimension() must match the Qdrant collection |
| Demo safety | ENVIRONMENT=demo will not boot unless registration is closed and DEMO_READONLY_EMAIL is set. That user cannot mutate; copilot and /v1/audit/export are allowlisted |
431 passing tests in pytest tests/unit with LLM keys blanked (CI job test). One more collected test is skipped; 37 requires_infra tests are a separate job.
Honest scope
Not in this project: model training, statistical inference, forecasting, a live Creem checkout, an observability stack, a mobile store build, or a live email/WhatsApp inbound relay (webhook secrets exist so unsigned requests can be rejected).
How to see it
- This page — use Request a walkthrough below to schedule a live session. Screenshots/recording after local capture (
docs/SCREENSHOTS.mdin the repo). - On request — seeded workspace, demo script: login → feed → inbox → copilot on
#8890→ timeline + export. Request a walkthrough. - Private repo — on request. Offline reproduce:
pip install -e ".[dev]"
python scripts/run_extraction_eval.py --split gold --predictor heuristic
python -m pytest tests/unit -q
Full stack: docker compose --profile app up -d --build, then scripts/seed_public_demo.py.
Highlights
- WhatsApp .txt and PDF ingest into a tenant-scoped knowledge graph, with a Bandeja review queue for the ambiguous cases (multi-topic or low-confidence routing)
- Hybrid retrieval: Qdrant vectors + FalkorDB commitment edges, fused by a lexical score boost, with sender-and-timestamp citations
- Entity F1 0.933 (0.923–0.944 over 3 LLM runs) vs 0.692 heuristic baseline on a 93-row hand-annotated corpus — 48 rows are deliberate negatives
- Every workspace-scoped route is membership-checked; a unit test walks the live FastAPI route table (431 passing offline tests)
- Commitment audit timeline and CSV/PDF export — who promised what, when, and when it changed
Challenges
Outcomes
- Ingest → chunk → extract → resolve → Qdrant + FalkorDB, with a seeded Colombian logistics workspace and audit-tier timelines
- Ops copilot (Spanish) with citation cards; delivery-promise questions naming an order number can short-circuit the LLM via deterministic synthesis; refusals when evidence is missing
- Extraction eval offline (`python scripts/run_extraction_eval.py --split gold --predictor heuristic`); 431 passing unit tests with LLM keys blanked in CI
- Web PWA is the shipping client (service worker + install prompt); Expo typechecks in CI but has no store build; Telegram bot code exists and is off on every billing tier
Live walkthrough
Lante has no public always-on URL — the stack is spun up on request. Get in touch, I'll approve access, we'll schedule a time, and I'll turn the project on for a live walkthrough.
Request a walkthrough