File size: 11,514 Bytes
2c8a4ec 329142e 2c8a4ec 6d96d0e 2c8a4ec ae1fdb3 2c8a4ec 68c07e1 329142e 6d96d0e f2930ce 72f78c8 2c8a4ec 70b1f2f 9671b99 70b1f2f 3c5341a 329142e 70b1f2f 9671b99 70b1f2f 9671b99 70b1f2f 329142e 70b1f2f 329142e 9671b99 1e6fa10 9671b99 329142e 70b1f2f 3c5341a 9671b99 3c5341a 9671b99 3c5341a 9671b99 70b1f2f 329142e 70b1f2f 329142e 70b1f2f 329142e 70b1f2f 329142e 1e6fa10 9671b99 1e6fa10 70b1f2f 329142e 70b1f2f 329142e 70b1f2f 329142e 70b1f2f 329142e 70b1f2f 329142e 70b1f2f 9671b99 70b1f2f 329142e 70b1f2f 6d290e0 70b1f2f 329142e 70b1f2f 329142e 70b1f2f 329142e 70b1f2f 329142e 70b1f2f 329142e 6ded0ab | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 | ---
title: Pay Equity For EU
emoji: 🐨
colorFrom: purple
colorTo: indigo
sdk: gradio
sdk_version: 6.18.0
python_version: '3.12'
app_file: main.py
pinned: false
license: apache-2.0
short_description: Know what you should earn
tags:
- build-small-hackathon
- backyard-ai
- gradio
- rag
- multilingual
- tiny-model
- tiny-titan
- off-brand
- nemotron
- pay-equity
- legal-tech
- denmark
- track:backyard
- sponsor:openbmb
- sponsor:nvidia
- achievement:offgrid
- achievement:offbrand
- achievement:sharing
- achievement:fieldnotes
---
# Pay Equity for EU
A user-friendly salary expectation tool for (currently) any danish engineers, to get an idea of what they deserve as a salary after suffering years of studies. Embark yourself in my app and see what salary you should be expecting in your future dream job, or even get to know what salary you actually should be having if you currently have a job! In parallel with an accessible salary overview, you can chat with a grounded, citation-first assistant that answers any questions you might have to prepare your next salary-talk or job-interview!
**"What pay should I expect?"** and **"What does the law actually let me ask for?"**
## Why I built this
I am a French and Danish engineer, and I built this for my engineering friends
and colleagues so they can find out whether they are actually being paid fairly,
or at least have the tools to know what they deserve after studying so hard. I
have been in that exact situation myself: not knowing what salary to expect, and
not knowing whether I had any reason to ask for a raise.
EU directives and laws are almost always overlooked by ordinary students,
because they never really reach us. The EU Pay Transparency Directive (the
"equal pay act") only entered into force very recently, on 6 June, which lands at
an ideal moment for this hackathon. I wanted to use that moment to help my fellow
engineers get the salary they deserve. Whether they are a man or a woman, they
deserve to know what they are worth.
Through this app I hope to give Danish engineers like me the proper tools to ask
for what they are worth, without worry or shame, and in the future to extend the
same help to economists, managers, political-science graduates, and to engineers
in France.
## The problem
In Denmark, salary negotiation is quite vague. New graduates rarely know
what the market pays, and how much they are actually worth compared to their peers. On top of that, the EU recently passed the **Pay Transparency
Directive (2023/970)**, which gives every employee the right to ask their employer for pay information, gender pay-gap reports, and more. But who understands what's written on these reports? So people
sign offers without knowing either the number or the rights they already have...
This app fixes both, in one chat and an accessible user-salary-creation (do not worry, you will not be disclosing anything private!)
## What it does
- **Expected pay** — deterministic lookup over parsed IDA *lønstatistik*
(salary statistics from the engineers' union), broken down by sector,
experience, and role. I use all this metadata to perform precise and effective salary chunk query.
- **Your rights** — RAG over the EU Pay Transparency Directive with
article-level citations on every claim (e.g. *"You can request the
average pay for your role broken down by gender (Art. 7(1))"*). Answers are
in **English** in v1; the directive text is already parsed in DA / FR and the
generation model is multilingual, so localized answers are an easy implementation!
- **Private document review** — session-only upload of a contract or payslip,
never persisted, never sent to a third-party API.
## Scope: Danish engineers first
To start with, this is deliberately focused on **Danish engineers**. The salary
side is built on the official salary statistics (*lønstatistik*) published by
**IDA** — *Ingeniørforeningen i Danmark*, the Danish engineers' trade union and
professional association (*fagforening* is simply Danish for "trade union"). IDA
surveys its members every year and publishes detailed pay tables by education,
sector, region, seniority, role, and gender. Those tables are exactly the
ground truth an engineer needs to answer "what should I earn?", so they are the
foundation here. I started with engineers since I am myself one and I know that salary conversations are always a pain.
## Future goals
There are two future goals which I would love to extend on:
### Extend to Djøf
**Djøf** is the other large Danish professional association and trade union, but
for a completely different group of people: those working in **law, economics,
business administration, political science, and the social sciences** (in
Denmark these members are often just called *djøfere*). Like IDA, Djøf publishes
its own annual salary statistics for its members. Because the salary lookup is
just a structured table behind a profile, adding Djøf is mostly a matter of
parsing and indexing their statistics into the same format. That would cover a
huge share of Danish white-collar graduates beyond engineering.
### Extend to France
I am also French, and the EU Pay Transparency Directive applies across the whole
European Union, not just Denmark. Each member state (including France) must turn
it into national law by June 2026, so the "what are my rights?" is also a fair question an frenchman could aks himself (that would then be "Quelles sont mes droits?!"). The next step there is to plug in French salary data sources so a French engineer gets the same profile-driven benchmark a Danish one does today.
## Hackathon track
**Backyard AI** — practical, problem-solving, runs on hardware you own, built
for a real circle of people (Danish graduates job-hunting right now).
## Tech stack
| Layer | Choice | Why |
|---|---|---|
| Generation LLM | **Cohere Tiny Aya 3B** | Multilingual (EN/DA/FR), under the Tiny Titan 4B cap |
| Swap-in LLM | OpenBMB MiniCPM3-4B | One-line swap via `LLMClient` interface |
| Embeddings | **NVIDIA Nemotron-Embed 1B v2** | Strong multilingual retrieval, ungated |
| Retrieval | FAISS IndexFlatIP + per-source top-k + Reciprocal Rank Fusion | Keeps salary numbers and legal text from drowning each other |
| Runtime | Gradio Server + ZeroGPU on HF Spaces | FastAPI-style `@app.api` endpoints behind a custom JS frontend |
| Frontend | **Hand-authored HTML / CSS / JS SPA** (no Gradio Blocks at the root) | Onboarding questionaire → profile-driven dashboard with hand-drawn SVG charts (bell curve, gauge, comparison bars) → contextual rights chat with seeded questions, served on top of `gradio.Server` via the JS client |
| Type / palette | Bricolage Grotesque + Inter on a sage / ink Scandinavian palette | Goes well beyond default Gradio look |
All model revisions are pinned by SHA in `configs/models.yaml`. No `latest`.
Total parameter footprint: **~4B** (Tiny Aya 3B + Nemotron-Embed 1B), well
under the 32B cap.
## Why small models matter here
Legal advice and salary numbers are exactly the domain where you don't want a
big confident model to free-associate. Grounding the small model on parsed
union tables and article-chunked directive text means every answer can be
checked against a citation. A small, transparent stack is the *right* tool for
"does this sentence come from the law or did the model make it up?"
## Demo video
[Watch the demo](https://youtu.be/TulJ8n-uR3A)
## Social post
[Read the post](https://www.linkedin.com/posts/%F0%9F%9A%80-simon-w-%C3%B8-larsen-957a97159_pay-equity-for-eu-build-small-hackathon-share-7472430584410091521-g_DP/?utm_source=share&utm_medium=member_desktop&rcm=ACoAACYR1ZMBGCnM42wuzlS3aS9FUP-eMPYGiz4)
## Run it locally
```bash
uv sync
uv run python main.py
```
Requires a built vector index at `data/processed/index/`. See
[`src/backend/README.md`](src/backend/README.md) for the RAG pipeline and
[`scripts/eval/`](scripts/eval/) for the offline benchmark.
## Disclaimer
This app surfaces public salary statistics and the published text of EU
Directive 2023/970. It is **not legal advice**. For binding interpretation,
talk to your union or a lawyer.
# Development log
### Jun 9 — Setup
Created the Hugging Face Space inside `build-small-hackathon`, joined the org, downloaded the EU Directive 2023/970 and the salary statistics PDF. Scoped v1 to Danish engineers under IDA (the engineer union) only; Djøf (white-collar jobs union) and France left as future work. Wired in ML-Intern, and Gradio Server as the core stack.
### Jun 10 — Knowledge extraction
Sent ML-Intern to parse both sources on HF Jobs (a10g-large GPU):
- **EU Directive** — Nemotron Parse over rasterized pages → 24 pages, 366 elements, JSON-Schema validated. Caught and removed the `TableInsertionLogitsProcessor`, which was forcing tabular output and mangling prose recitals into fake tables.
- **IDA + Djøf lønstatistik** — separate extractors (lattice for IDA, header-anchored for Djøf's borderless tables), normalized into typed records with derived experience years and citation-bearing RAG text. 266 non-salary rows dropped during cleanup.
Both outputs published as private HF Datasets.
### Jun 14 — Index and retrieval
Embedded all chunks on Modal (local CPU too slow for the Nemotron 1B encoder), assembled a FAISS `IndexFlatIP` locally. Hit HF's 10MB non-LFS file ceiling — installed git-lfs and shipped the 20MB `index.faiss` as a pointer.
### Jun 15 — RAG, frontend, chat
Four major pieces in one push:
1. **Dual-source Reciprocal Rank Fusion** — query the directive and IDA independently, fuse the ranked lists.
2. **Typed salary layer** — `IdaSalaryRecord` schema, `Profile` / `Dashboard` models, wildcard matching, specificity ranking, reliability-tier gating, 2026 temporal projection.
3. **Re-extracted IDA tables** with a panel-aware `pdfplumber` parser, recovering multi-axis data (region, gender × leadership, public position) that v1 had flattened — 559 records. This restored the gender pay-gap data, the core of the equal-pay feature. Biggest data win of the project.
4. **Custom frontend** — hand-authored HTML/CSS/JS SPA served by `gradio.Server` (no Gradio Blocks at the root), profile-driven dashboard with hand-drawn SVG visuals, Branch A (employed) and Branch B (job-seeking) onboarding flows, contextual rights chat. Sage/ink Scandinavian palette.
GPU warmup triggered on the first wizard answer so Tiny Aya is ready by the time the user reaches the chat. Added token streaming, then removed it for a cleaner single response.
### Jun 15–16 — The number bug
Salaries rendering ~10× too small. The small multilingual LLM was reading thousands separators as decimal points: first the Danish `47.010` → `4,701`, then the en-US `47,300` → `4,730` after switching format. Fix: bare integers with no separators in everything the model reads (`rag_text`, injected context), plus a "copy digits exactly" system instruction. Regenerated `chunks.jsonl` only — FAISS rows stayed aligned, so no re-embed needed.
### Jun 16 — Submission
Scrubbed a write token that had leaked into the git remote URL. Final README pass — motivation, IDA scope, Djøf and France as future work, track/sponsor/achievement tags, demo video, social links. Added a "coming soon: Djøf" dashboard banner. Last issue was a ZeroGPU container prestart-hook error on the Space — diagnosed as platform infrastructure (no code changes resolved it), fixed by a Factory reboot. |