Franskaman110's picture
Frontend
1e6fa10
|
Raw History Blame Contribute Delete
4.74 kB

A newer version of the Gradio SDK is available: 6.30.0

Upgrade

Backend

A retrieval-augmented chat assistant over two sources — IDA salary statistics (engineers) and the EU Pay Transparency Directive 2023/970 — served through a Gradio Server. build_server() (in server.py) creates the gradio.Server, registers the backend API endpoints, and serves the custom HTML/CSS/JS SPA from ../frontend/ at /. The SPA calls the registered @app.api endpoints via the Gradio JS client.

Start it

From the repo root:

uv run python main.py

main.py calls build_server().launch(...). The server prints a local URL (default http://127.0.0.1:7860). Open it to use the SPA (wizard → dashboard → rights chat).

Before first run, build the vector index (see RAG pipeline below).

RAG pipeline

query ─► embed (Nemotron) ─► per-source top-k ─► RRF fuse ─► grounded prompt ─► Tiny Aya ─► cited answer
  • Ingestion (indexing/ingest.py): the Directive document.json is grouped into one Chunk per article (Art. N citations + metadata.article_number / article_title); IDA rows (union == "IDA") become chunks from their ready-made rag_text (cited by source page) with structured salary metadata (sector, category, experience, measures, …). Djøf is excluded for now. Only Chunk.text is embedded; metadata is persisted in chunks.jsonl for downstream use.
  • Offline index build (scripts/build_index.py): embeds the whole corpus on ZeroGPU (the embedding call is wrapped in @spaces.GPU) and persists data/processed/index/{index.faiss, chunks.jsonl}. Run once on the ZeroGPU Space (or any GPU machine) and commit the output — the live app only loads the index and never re-embeds the corpus at runtime.
  • Metadata refresh (scripts/refresh_metadata.py): rewrites chunks.jsonl from updated ingest logic without re-embedding. Safe only when chunk ids and order are unchanged; otherwise re-run build_index.py.
  • Runtime (rag.py): answer_stream() is the single grounding path shared by the UI and (via answer()) the /chat API. Retrieval runs per source (directive, lonstatistik) — top-5 each — then Reciprocal Rank Fusion into the final context (default 6 chunks). Query embedding + generation run inside one @spaces.GPU function (_retrieve_and_stream).

Models are pinned by revision in configs/models.yaml. The LLM backend is selected by llm.<role>.backend (tiny_aya | minicpm | llama_cpp | echo); the embedder by embeddings.repo + loader.

Endpoints

Registered in api/endpoints.py via @app.api(name=...) and consumed by the custom SPA in ../frontend/static/app.js through the Gradio JS client.

api_name Input Returns Status
/chat messages: list[dict], lang: str ChatResponse (reply + citations) RAG-grounded
/dashboard profile: Profile Dashboard (KPIs + cells + projection) profile-driven IDA lookup
/parse_directive lang: str list[Chunk] real (article chunks)
/parse_lonstatistik file: FileData, source: str ParseResult placeholder
/index_documents source: str IndexStatus reports loaded index

Call the chat endpoint (smoke test)

from gradio_client import Client

c = Client("http://127.0.0.1:7860")
r = c.predict(
    messages=[{"role": "user", "content": "What information am I entitled to about pay?"}],
    lang="en", api_name="/chat",
)
print(r["reply"], r["citations"])

Layout

server.py            build_server() — Server + endpoints + serves the custom SPA at /
rag.py               answer() + single @spaces.GPU retrieve→generate function
config.py            load_models_config() (configs/models.yaml) + get_secret()
schemas.py           pydantic models (Chunk, ChatMessage/Request/Response, …)
llm/client.py        LLMClient ABC, TinyAyaClient, MiniCPMClient, LlamaCppClient(stub), get_llm_client()
llm/prompt.py        build_grounded_prompt(query, chunks, lang)
indexing/ingest.py   directive + IDA → list[Chunk]
indexing/embedder.py Embedder (config-driven, sentence-transformers | transformers)
indexing/store.py    VectorStore (FAISS IndexFlatIP + persist/load/search)
parsing/             directive.py (→ ingest), lonstatistik.py (placeholder)
api/endpoints.py     register(app): the @app.api endpoints
scripts/build_index.py       offline corpus embedding → data/processed/index/
scripts/refresh_metadata.py  rewrite chunks.jsonl metadata without re-embedding
scripts/eval/run_benchmark.py  offline retrieval benchmark (see root README.md)