Download src/backend/README.md from build-small-hackathon/pay-equity-for-eu: direct link, hf CLI and curl.
- Browser
- Download file 4.74 kB
-
https://huggingface.co/spaces/build-small-hackathon/pay-equity-for-eu/resolve/main/src/backend/README.md
- Command line
-
hf download hf://spaces/build-small-hackathon/pay-equity-for-eu/src/backend/README.md
-
curl -L -o README.md https://huggingface.co/spaces/build-small-hackathon/pay-equity-for-eu/resolve/main/src/backend/README.md
A newer version of the Gradio SDK is available: 6.30.0
Backend
A retrieval-augmented chat assistant over two sources — IDA salary
statistics (engineers) and the EU Pay Transparency Directive 2023/970 —
served through a Gradio Server.
build_server() (in server.py) creates the gradio.Server, registers the
backend API endpoints, and serves the custom HTML/CSS/JS SPA from
../frontend/ at /. The SPA calls the registered
@app.api endpoints via the Gradio JS client.
Start it
From the repo root:
uv run python main.py
main.py calls build_server().launch(...). The server prints a local URL
(default http://127.0.0.1:7860). Open it to use the SPA (wizard → dashboard
→ rights chat).
Before first run, build the vector index (see RAG pipeline below).
RAG pipeline
query ─► embed (Nemotron) ─► per-source top-k ─► RRF fuse ─► grounded prompt ─► Tiny Aya ─► cited answer
- Ingestion (
indexing/ingest.py): the Directivedocument.jsonis grouped into oneChunkper article (Art. Ncitations +metadata.article_number/article_title); IDA rows (union == "IDA") become chunks from their ready-maderag_text(cited by source page) with structured salarymetadata(sector, category, experience, measures, …). Djøf is excluded for now. OnlyChunk.textis embedded;metadatais persisted inchunks.jsonlfor downstream use. - Offline index build (
scripts/build_index.py): embeds the whole corpus on ZeroGPU (the embedding call is wrapped in@spaces.GPU) and persistsdata/processed/index/{index.faiss, chunks.jsonl}. Run once on the ZeroGPU Space (or any GPU machine) and commit the output — the live app only loads the index and never re-embeds the corpus at runtime. - Metadata refresh (
scripts/refresh_metadata.py): rewriteschunks.jsonlfrom updated ingest logic without re-embedding. Safe only when chunk ids and order are unchanged; otherwise re-runbuild_index.py. - Runtime (
rag.py):answer_stream()is the single grounding path shared by the UI and (viaanswer()) the/chatAPI. Retrieval runs per source (directive,lonstatistik) — top-5 each — then Reciprocal Rank Fusion into the final context (default 6 chunks). Query embedding + generation run inside one@spaces.GPUfunction (_retrieve_and_stream).
Models are pinned by revision in configs/models.yaml. The LLM backend is
selected by llm.<role>.backend (tiny_aya | minicpm | llama_cpp |
echo); the embedder by embeddings.repo + loader.
Endpoints
Registered in api/endpoints.py via @app.api(name=...) and consumed by the
custom SPA in ../frontend/static/app.js through the Gradio JS client.
api_name |
Input | Returns | Status |
|---|---|---|---|
/chat |
messages: list[dict], lang: str |
ChatResponse (reply + citations) |
RAG-grounded |
/dashboard |
profile: Profile |
Dashboard (KPIs + cells + projection) |
profile-driven IDA lookup |
/parse_directive |
lang: str |
list[Chunk] |
real (article chunks) |
/parse_lonstatistik |
file: FileData, source: str |
ParseResult |
placeholder |
/index_documents |
source: str |
IndexStatus |
reports loaded index |
Call the chat endpoint (smoke test)
from gradio_client import Client
c = Client("http://127.0.0.1:7860")
r = c.predict(
messages=[{"role": "user", "content": "What information am I entitled to about pay?"}],
lang="en", api_name="/chat",
)
print(r["reply"], r["citations"])
Layout
server.py build_server() — Server + endpoints + serves the custom SPA at /
rag.py answer() + single @spaces.GPU retrieve→generate function
config.py load_models_config() (configs/models.yaml) + get_secret()
schemas.py pydantic models (Chunk, ChatMessage/Request/Response, …)
llm/client.py LLMClient ABC, TinyAyaClient, MiniCPMClient, LlamaCppClient(stub), get_llm_client()
llm/prompt.py build_grounded_prompt(query, chunks, lang)
indexing/ingest.py directive + IDA → list[Chunk]
indexing/embedder.py Embedder (config-driven, sentence-transformers | transformers)
indexing/store.py VectorStore (FAISS IndexFlatIP + persist/load/search)
parsing/ directive.py (→ ingest), lonstatistik.py (placeholder)
api/endpoints.py register(app): the @app.api endpoints
scripts/build_index.py offline corpus embedding → data/processed/index/
scripts/refresh_metadata.py rewrite chunks.jsonl metadata without re-embedding
scripts/eval/run_benchmark.py offline retrieval benchmark (see root README.md)