Download README.md from build-small-hackathon/pay-equity-for-eu: direct link, hf CLI and curl.
- Browser
- Download file 11.5 kB
-
https://huggingface.co/spaces/build-small-hackathon/pay-equity-for-eu/resolve/main/README.md
- Command line
-
hf download hf://spaces/build-small-hackathon/pay-equity-for-eu/README.md
-
curl -L -o README.md https://huggingface.co/spaces/build-small-hackathon/pay-equity-for-eu/resolve/main/README.md
A newer version of the Gradio SDK is available: 6.30.0
title: Pay Equity For EU
emoji: 🐨
colorFrom: purple
colorTo: indigo
sdk: gradio
sdk_version: 6.18.0
python_version: '3.12'
app_file: main.py
pinned: false
license: apache-2.0
short_description: Know what you should earn
tags:
- build-small-hackathon
- backyard-ai
- gradio
- rag
- multilingual
- tiny-model
- tiny-titan
- off-brand
- nemotron
- pay-equity
- legal-tech
- denmark
- track:backyard
- sponsor:openbmb
- sponsor:nvidia
- achievement:offgrid
- achievement:offbrand
- achievement:sharing
- achievement:fieldnotes
Pay Equity for EU
A user-friendly salary expectation tool for (currently) any danish engineers, to get an idea of what they deserve as a salary after suffering years of studies. Embark yourself in my app and see what salary you should be expecting in your future dream job, or even get to know what salary you actually should be having if you currently have a job! In parallel with an accessible salary overview, you can chat with a grounded, citation-first assistant that answers any questions you might have to prepare your next salary-talk or job-interview!
"What pay should I expect?" and "What does the law actually let me ask for?"
Why I built this
I am a French and Danish engineer, and I built this for my engineering friends and colleagues so they can find out whether they are actually being paid fairly, or at least have the tools to know what they deserve after studying so hard. I have been in that exact situation myself: not knowing what salary to expect, and not knowing whether I had any reason to ask for a raise.
EU directives and laws are almost always overlooked by ordinary students, because they never really reach us. The EU Pay Transparency Directive (the "equal pay act") only entered into force very recently, on 6 June, which lands at an ideal moment for this hackathon. I wanted to use that moment to help my fellow engineers get the salary they deserve. Whether they are a man or a woman, they deserve to know what they are worth.
Through this app I hope to give Danish engineers like me the proper tools to ask for what they are worth, without worry or shame, and in the future to extend the same help to economists, managers, political-science graduates, and to engineers in France.
The problem
In Denmark, salary negotiation is quite vague. New graduates rarely know what the market pays, and how much they are actually worth compared to their peers. On top of that, the EU recently passed the Pay Transparency Directive (2023/970), which gives every employee the right to ask their employer for pay information, gender pay-gap reports, and more. But who understands what's written on these reports? So people sign offers without knowing either the number or the rights they already have...
This app fixes both, in one chat and an accessible user-salary-creation (do not worry, you will not be disclosing anything private!)
What it does
- Expected pay — deterministic lookup over parsed IDA lønstatistik (salary statistics from the engineers' union), broken down by sector, experience, and role. I use all this metadata to perform precise and effective salary chunk query.
- Your rights — RAG over the EU Pay Transparency Directive with article-level citations on every claim (e.g. "You can request the average pay for your role broken down by gender (Art. 7(1))"). Answers are in English in v1; the directive text is already parsed in DA / FR and the generation model is multilingual, so localized answers are an easy implementation!
- Private document review — session-only upload of a contract or payslip, never persisted, never sent to a third-party API.
Scope: Danish engineers first
To start with, this is deliberately focused on Danish engineers. The salary side is built on the official salary statistics (lønstatistik) published by IDA — Ingeniørforeningen i Danmark, the Danish engineers' trade union and professional association (fagforening is simply Danish for "trade union"). IDA surveys its members every year and publishes detailed pay tables by education, sector, region, seniority, role, and gender. Those tables are exactly the ground truth an engineer needs to answer "what should I earn?", so they are the foundation here. I started with engineers since I am myself one and I know that salary conversations are always a pain.
Future goals
There are two future goals which I would love to extend on:
Extend to Djøf
Djøf is the other large Danish professional association and trade union, but for a completely different group of people: those working in law, economics, business administration, political science, and the social sciences (in Denmark these members are often just called djøfere). Like IDA, Djøf publishes its own annual salary statistics for its members. Because the salary lookup is just a structured table behind a profile, adding Djøf is mostly a matter of parsing and indexing their statistics into the same format. That would cover a huge share of Danish white-collar graduates beyond engineering.
Extend to France
I am also French, and the EU Pay Transparency Directive applies across the whole European Union, not just Denmark. Each member state (including France) must turn it into national law by June 2026, so the "what are my rights?" is also a fair question an frenchman could aks himself (that would then be "Quelles sont mes droits?!"). The next step there is to plug in French salary data sources so a French engineer gets the same profile-driven benchmark a Danish one does today.
Hackathon track
Backyard AI — practical, problem-solving, runs on hardware you own, built for a real circle of people (Danish graduates job-hunting right now).
Tech stack
| Layer | Choice | Why |
|---|---|---|
| Generation LLM | Cohere Tiny Aya 3B | Multilingual (EN/DA/FR), under the Tiny Titan 4B cap |
| Swap-in LLM | OpenBMB MiniCPM3-4B | One-line swap via LLMClient interface |
| Embeddings | NVIDIA Nemotron-Embed 1B v2 | Strong multilingual retrieval, ungated |
| Retrieval | FAISS IndexFlatIP + per-source top-k + Reciprocal Rank Fusion | Keeps salary numbers and legal text from drowning each other |
| Runtime | Gradio Server + ZeroGPU on HF Spaces | FastAPI-style @app.api endpoints behind a custom JS frontend |
| Frontend | Hand-authored HTML / CSS / JS SPA (no Gradio Blocks at the root) | Onboarding questionaire → profile-driven dashboard with hand-drawn SVG charts (bell curve, gauge, comparison bars) → contextual rights chat with seeded questions, served on top of gradio.Server via the JS client |
| Type / palette | Bricolage Grotesque + Inter on a sage / ink Scandinavian palette | Goes well beyond default Gradio look |
All model revisions are pinned by SHA in configs/models.yaml. No latest.
Total parameter footprint: ~4B (Tiny Aya 3B + Nemotron-Embed 1B), well under the 32B cap.
Why small models matter here
Legal advice and salary numbers are exactly the domain where you don't want a big confident model to free-associate. Grounding the small model on parsed union tables and article-chunked directive text means every answer can be checked against a citation. A small, transparent stack is the right tool for "does this sentence come from the law or did the model make it up?"
Demo video
Social post
Run it locally
uv sync
uv run python main.py
Requires a built vector index at data/processed/index/. See
src/backend/README.md for the RAG pipeline and
scripts/eval/ for the offline benchmark.
Disclaimer
This app surfaces public salary statistics and the published text of EU Directive 2023/970. It is not legal advice. For binding interpretation, talk to your union or a lawyer.
Development log
Jun 9 — Setup
Created the Hugging Face Space inside build-small-hackathon, joined the org, downloaded the EU Directive 2023/970 and the salary statistics PDF. Scoped v1 to Danish engineers under IDA (the engineer union) only; Djøf (white-collar jobs union) and France left as future work. Wired in ML-Intern, and Gradio Server as the core stack.
Jun 10 — Knowledge extraction
Sent ML-Intern to parse both sources on HF Jobs (a10g-large GPU):
- EU Directive — Nemotron Parse over rasterized pages → 24 pages, 366 elements, JSON-Schema validated. Caught and removed the
TableInsertionLogitsProcessor, which was forcing tabular output and mangling prose recitals into fake tables. - IDA + Djøf lønstatistik — separate extractors (lattice for IDA, header-anchored for Djøf's borderless tables), normalized into typed records with derived experience years and citation-bearing RAG text. 266 non-salary rows dropped during cleanup.
Both outputs published as private HF Datasets.
Jun 14 — Index and retrieval
Embedded all chunks on Modal (local CPU too slow for the Nemotron 1B encoder), assembled a FAISS IndexFlatIP locally. Hit HF's 10MB non-LFS file ceiling — installed git-lfs and shipped the 20MB index.faiss as a pointer.
Jun 15 — RAG, frontend, chat
Four major pieces in one push:
- Dual-source Reciprocal Rank Fusion — query the directive and IDA independently, fuse the ranked lists.
- Typed salary layer —
IdaSalaryRecordschema,Profile/Dashboardmodels, wildcard matching, specificity ranking, reliability-tier gating, 2026 temporal projection. - Re-extracted IDA tables with a panel-aware
pdfplumberparser, recovering multi-axis data (region, gender × leadership, public position) that v1 had flattened — 559 records. This restored the gender pay-gap data, the core of the equal-pay feature. Biggest data win of the project. - Custom frontend — hand-authored HTML/CSS/JS SPA served by
gradio.Server(no Gradio Blocks at the root), profile-driven dashboard with hand-drawn SVG visuals, Branch A (employed) and Branch B (job-seeking) onboarding flows, contextual rights chat. Sage/ink Scandinavian palette.
GPU warmup triggered on the first wizard answer so Tiny Aya is ready by the time the user reaches the chat. Added token streaming, then removed it for a cleaner single response.
Jun 15–16 — The number bug
Salaries rendering ~10× too small. The small multilingual LLM was reading thousands separators as decimal points: first the Danish 47.010 → 4,701, then the en-US 47,300 → 4,730 after switching format. Fix: bare integers with no separators in everything the model reads (rag_text, injected context), plus a "copy digits exactly" system instruction. Regenerated chunks.jsonl only — FAISS rows stayed aligned, so no re-embed needed.
Jun 16 — Submission
Scrubbed a write token that had leaked into the git remote URL. Final README pass — motivation, IDA scope, Djøf and France as future work, track/sponsor/achievement tags, demo video, social links. Added a "coming soon: Djøf" dashboard banner. Last issue was a ZeroGPU container prestart-hook error on the Space — diagnosed as platform infrastructure (no code changes resolved it), fixed by a Factory reboot.