Spaces:
Sleeping
Sleeping
Commit ·
1ec12ef
1
Parent(s): 9badb7a
feat(phase-11): kill JD-scraping leak in keyword pool — V2-only
Browse filesScoped to V2 per request (plan assumed shared V1+V2 chokepoint). New
filter_scraped_noise()/_is_skill_like() in src/external_ats.py, applied only in
src/resume_v2_natural.py::_allocate_keywords. Rejects geography (graphql safe from
'hq'), the hiring company's own name, cue-gated person names (so job-title
headings survive), corporate-entity/role nouns, and job-board UI blurbs. V1
extractor path untouched. New tests/test_jd_leak_filter.py (6 tests); V1 (11) + V2
(17) suites green — 34 total.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- .planning/REQUIREMENTS.md +3 -0
- .planning/phases/11-jd-scraping-leak-fix/11-01-SUMMARY.md +61 -0
- .planning/phases/11-jd-scraping-leak-fix/11-CONTEXT.md +78 -0
- HISTORY.md +33 -0
- README.md +11 -0
- src/external_ats.py +143 -0
- src/resume_v2_natural.py +10 -4
- tests/test_jd_leak_filter.py +175 -0
.planning/REQUIREMENTS.md
CHANGED
|
@@ -106,3 +106,6 @@ V2 must NOT append keyword lists. Instead, fan out across the existing model poo
|
|
| 106 |
|
| 107 |
## R26: Label V1 + documentation + default-safe
|
| 108 |
Explicitly label the current behavior as "V1" in README.md and HISTORY.md, document V2, and ensure V1 stays the default so existing flows are unaffected if no version is specified.
|
|
|
|
|
|
|
|
|
|
|
|
| 106 |
|
| 107 |
## R26: Label V1 + documentation + default-safe
|
| 108 |
Explicitly label the current behavior as "V1" in README.md and HISTORY.md, document V2, and ensure V1 stays the default so existing flows are unaffected if no version is specified.
|
| 109 |
+
|
| 110 |
+
## R27: Kill the JD-scraping leak in the keyword extractor
|
| 111 |
+
The keyword extraction layer (the code that turns a raw JD into the keyword pool consumed by BOTH V1 structured placement and V2 sentence generation) is leaking scraped, non-skill text directly into the candidate's resume. Real evidence from a generated Backbase resume: the Skills "Other" block contained the company CEO's name (`jouk pleiter`), geography boilerplate (`headquartered, amsterdam, north america europe, middle east asia-pacific africa, latin america`), About-us marketing copy (`backbase helped financial institutions … bold climate action … worldwide network, backbasers partners`), and recruiter-blurb fragments (`re interested, likely, recruiter, don, show`); the injected summary contained scraped LinkedIn UI text (`actively engaged on LinkedIn Corporation to get alerts, reach applicants, and stay a top applicant`) and a stray collaborator name (`Mohammed Nayeem`). The extractor MUST NOT emit: company marketing/boilerplate copy, person/named-entity tokens (executives, recruiters, collaborators), location/geography lists, multi-word sentence fragments, or generic recruiter/UI-blurb tokens. Only genuine, role-relevant skill / tool / domain keywords survive. Scope is the extraction + filtering layer ONLY — NOT the honesty guardrails, injection grammar, or fit-gate (deferred). Success: a JD containing company boilerplate + a CEO name + a geography list yields a keyword pool with zero of those tokens, and a regenerated resume's Skills "Other" line and summary contain no scraped non-skill text.
|
.planning/phases/11-jd-scraping-leak-fix/11-01-SUMMARY.md
ADDED
|
@@ -0,0 +1,61 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
phase: 11-jd-scraping-leak-fix
|
| 3 |
+
plan: 01
|
| 4 |
+
type: summary
|
| 5 |
+
status: DONE
|
| 6 |
+
requirements: [R27]
|
| 7 |
+
deviation: "Scoped to V2-only per user instruction (plan assumed shared V1+V2 chokepoint)"
|
| 8 |
+
---
|
| 9 |
+
|
| 10 |
+
# 11-01 Summary: Kill the JD-Scraping Leak (V2-only)
|
| 11 |
+
|
| 12 |
+
## Deviation from plan
|
| 13 |
+
The verified plan fixed the **shared** chokepoint (`_is_term_like` in
|
| 14 |
+
`external_ats.py`), which would have changed both V1 and V2. The user instructed
|
| 15 |
+
"ensure this only goes in V2." Implementation was re-scoped: the filter is a
|
| 16 |
+
standalone function applied **only in the V2 allocation path**; the shared V1
|
| 17 |
+
extractor is untouched.
|
| 18 |
+
|
| 19 |
+
## What was built
|
| 20 |
+
|
| 21 |
+
### src/external_ats.py (additive — definitions only, not wired into V1)
|
| 22 |
+
- `_GEO_EXTENDED`, `_ENTITY_NOISE`, `_ROLE_NOISE`, `_SENTENCE_STARTERS`,
|
| 23 |
+
`_UI_BLURB`, `_NAME_CUE_BEFORE`
|
| 24 |
+
- `_contains_geo_token()` — word-boundary single-word match (so `hq` never
|
| 25 |
+
blocks `graphql`) + substring match for multi-word geo phrases
|
| 26 |
+
- `_contains_person_name()` — **cue-gated**: a Title-Case bigram is a name only
|
| 27 |
+
if it sits next to a person cue ("Recruiter: …", "… is the CEO", "— Manager").
|
| 28 |
+
Scans all adjacent pairs, so a name embedded in a longer gram is caught. Avoids
|
| 29 |
+
false-flagging job-title/domain headings ("Forward Deployed", "Banking
|
| 30 |
+
Integrations", "Product Manager")
|
| 31 |
+
- `_is_skill_like(term, original_jd, company)` — the rejection gate: UI noise,
|
| 32 |
+
UI blurbs, geography, role/entity nouns, the hiring company's own name
|
| 33 |
+
(word-boundary + s/rs/ers suffix), and cue-gated person names
|
| 34 |
+
- `filter_scraped_noise(terms, original_jd, company)` — public V2 entry point
|
| 35 |
+
|
| 36 |
+
### src/resume_v2_natural.py (V2 path only)
|
| 37 |
+
- `_allocate_keywords()` gained a `company` param; calls
|
| 38 |
+
`filter_scraped_noise(decision["includable"], jd_text, company)` right after
|
| 39 |
+
`decide_includable_terms`
|
| 40 |
+
- `generate_v2()` passes `company` into `_allocate_keywords`
|
| 41 |
+
|
| 42 |
+
### tests/test_jd_leak_filter.py (new, 6 tests)
|
| 43 |
+
- `test_bad_tokens_absent`, `test_good_tokens_present`, `test_no_false_positives`
|
| 44 |
+
- `test_v2_allocation_excludes_scraped` — end-to-end via `_allocate_keywords`
|
| 45 |
+
- `test_v1_extractor_unchanged` — **proves V2-only**: `amsterdam` still in the V1
|
| 46 |
+
`extract_external_keywords` pool, removed only by `filter_scraped_noise`
|
| 47 |
+
- `test_real_jd_coverage_not_regressed` — clean JD keeps ≥90% of its pool
|
| 48 |
+
|
| 49 |
+
### Docs
|
| 50 |
+
- README "Generation Modes" → V2 scraped-noise filter subsection (notes V1 not filtered)
|
| 51 |
+
- HISTORY → Phase 11 entry
|
| 52 |
+
|
| 53 |
+
## Verification
|
| 54 |
+
- `pytest tests/test_jd_leak_filter.py tests/test_structured_placement.py tests/test_resume_v2.py -q` → **34 passed**
|
| 55 |
+
- Real Backbase JD, V2 pool after fix: `['banking integrations api management',
|
| 56 |
+
'forward deployed product manager', 'gcp', 'ipaas reconciliation agile',
|
| 57 |
+
'rest soap graphql webhooks']` — zero names/geography/company/boilerplate.
|
| 58 |
+
|
| 59 |
+
## Out of scope (deferred, unchanged)
|
| 60 |
+
Honesty guardrails (`_REGULATED_CRED`/`_specialty_hit`), injection grammar/
|
| 61 |
+
coherence, fit-gate, upstream scraper. V1 placement behavior.
|
.planning/phases/11-jd-scraping-leak-fix/11-CONTEXT.md
ADDED
|
@@ -0,0 +1,78 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
phase: 11-jd-scraping-leak-fix
|
| 3 |
+
type: context
|
| 4 |
+
captured: 2026-06-25
|
| 5 |
+
source: real generated resume (Backbase Forward Deployed PM) + user-confirmed scope
|
| 6 |
+
---
|
| 7 |
+
|
| 8 |
+
# Phase 11 Context — Kill the JD-Scraping Leak
|
| 9 |
+
|
| 10 |
+
## Origin
|
| 11 |
+
A live resume generated for a **Backbase "Forward Deployed Product Manager for
|
| 12 |
+
Banking Integrations"** role (LinkedIn job 4415276270) was analyzed as if by the
|
| 13 |
+
hiring HR head. It would be rejected instantly — not for tuning reasons but
|
| 14 |
+
because the **keyword extractor leaked raw scraped JD/company text directly into
|
| 15 |
+
the resume**, producing fabrication-looking content and obvious automated-stuffing
|
| 16 |
+
signals.
|
| 17 |
+
|
| 18 |
+
## The actual leaked tokens (verbatim evidence)
|
| 19 |
+
From `Saiteja_Tirunagari_Backbase_Resume.tex`:
|
| 20 |
+
|
| 21 |
+
1. **Skills "Other" block** (`% ats-skills-other`, line 266) contained:
|
| 22 |
+
- Executive name: `jouk pleiter` (Backbase's CEO)
|
| 23 |
+
- Geography list: `headquartered, amsterdam, north america europe, middle east asia-pacific africa, latin america`
|
| 24 |
+
- About-us marketing copy: `backbase helped financial institutions, institutions stand, re taking bold climate, action helping future-proof, planet alongside, worldwide network, backbasers partners, peers, members, share`
|
| 25 |
+
- Recruiter/UI-blurb fragments: `re interested, likely, recruiter, don, show`
|
| 26 |
+
- Other sentence fragments: `things grand central, voice, client inside`
|
| 27 |
+
|
| 28 |
+
2. **Injected summary** (`% ats-injected-start`, line 148) contained:
|
| 29 |
+
- Scraped LinkedIn UI text: `actively engaged on LinkedIn Corporation to get alerts, reach applicants, and stay a top applicant in competitive PM roles`
|
| 30 |
+
- Unsupported claim tokens: `GCP, banking integrations`
|
| 31 |
+
|
| 32 |
+
3. **Injected bullets** referenced a stray collaborator name: `Mohammed Nayeem`
|
| 33 |
+
(bleed-through from unrelated profile/JD data).
|
| 34 |
+
|
| 35 |
+
## Root-cause hypothesis (to confirm in research)
|
| 36 |
+
The JD→keyword extraction layer is treating the ENTIRE scraped page (company
|
| 37 |
+
"About us", CEO bio, recruiter blurb, LinkedIn chrome) as JD body, then tokenizing
|
| 38 |
+
it into the keyword pool with no named-entity / boilerplate / fragment filtering.
|
| 39 |
+
This pool feeds BOTH:
|
| 40 |
+
- V1 structured placement (`src/latex_resume.py` — `decide_includable_terms`,
|
| 41 |
+
inject into `% ats-skills-other` + `% ats-injected-start`)
|
| 42 |
+
- V2 sentence generation (`src/resume_v2_natural.py` — `_allocate_keywords`)
|
| 43 |
+
|
| 44 |
+
## Scope decision (USER-CONFIRMED)
|
| 45 |
+
**IN scope:** the extraction + filtering layer only — stop non-skill tokens from
|
| 46 |
+
ever entering the keyword pool.
|
| 47 |
+
|
| 48 |
+
The extractor MUST NOT emit:
|
| 49 |
+
- company marketing / boilerplate copy
|
| 50 |
+
- person / named-entity tokens (executives, recruiters, collaborators)
|
| 51 |
+
- location / geography lists
|
| 52 |
+
- multi-word sentence fragments
|
| 53 |
+
- generic recruiter / UI-blurb tokens
|
| 54 |
+
|
| 55 |
+
Only genuine, role-relevant **skill / tool / domain** keywords survive.
|
| 56 |
+
|
| 57 |
+
**OUT of scope (explicitly deferred by the user — do NOT touch):**
|
| 58 |
+
- honesty guardrails (`_REGULATED_CRED`, `_specialty_hit`)
|
| 59 |
+
- grammar/coherence of the injected sentences
|
| 60 |
+
- fit-gate / domain-mismatch down-ranking
|
| 61 |
+
- the upstream JD scraper itself (we filter at extraction, not at scrape)
|
| 62 |
+
|
| 63 |
+
## Constraints
|
| 64 |
+
- Fix at the SHARED extraction layer so both V1 and V2 benefit from one change.
|
| 65 |
+
- Do NOT over-filter: legitimate keywords (REST, SOAP, GraphQL, iPaaS, Agile,
|
| 66 |
+
banking-integration domain terms) must still survive — no regression in real
|
| 67 |
+
keyword coverage.
|
| 68 |
+
- Existing tests must stay green: `tests/test_structured_placement.py` (V1, 11
|
| 69 |
+
tests), `tests/test_resume_v2.py` (V2, 17 tests).
|
| 70 |
+
- Project rule [[feedback_docs_update]]: update README.md + HISTORY.md.
|
| 71 |
+
- Honesty principle [[ats_no_stuffing_principle]]: this is the mechanism behind
|
| 72 |
+
"keyword dumps get stripped" — fixing extraction is the upstream cure.
|
| 73 |
+
|
| 74 |
+
## Success (acceptance)
|
| 75 |
+
A JD containing company boilerplate + a CEO name + a geography list yields a
|
| 76 |
+
keyword pool with **zero** of those tokens; a regenerated resume's Skills "Other"
|
| 77 |
+
line and summary contain no scraped non-skill text; true-keyword coverage does
|
| 78 |
+
not regress; both V1 and V2 test suites pass.
|
HISTORY.md
CHANGED
|
@@ -4,6 +4,39 @@ A running log of everything built, fixed, and changed. Most recent first.
|
|
| 4 |
|
| 5 |
---
|
| 6 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 7 |
## 2026-06-25 — Designation + Filename Update
|
| 8 |
|
| 9 |
- **Designation**: Changed NxtWave role from "Internal Product Manager" to
|
|
|
|
| 4 |
|
| 5 |
---
|
| 6 |
|
| 7 |
+
## 2026-06-25 — Phase 11: Kill the JD-Scraping Leak (V2-only)
|
| 8 |
+
|
| 9 |
+
A live Backbase resume exposed the keyword extractor leaking scraped, non-skill
|
| 10 |
+
text into the resume: the Skills "Other" block held the CEO's name (`jouk
|
| 11 |
+
pleiter`), geography (`amsterdam`, `latin america`), About-us copy (`worldwide
|
| 12 |
+
network`, `backbasers partners`), and recruiter fragments; the summary held
|
| 13 |
+
LinkedIn UI chrome (`actively engaged on LinkedIn Corporation…`) and a stray
|
| 14 |
+
collaborator name (`mohammed nayeem`).
|
| 15 |
+
|
| 16 |
+
- **Fix (scoped to V2 only, per request):** new `filter_scraped_noise()` +
|
| 17 |
+
`_is_skill_like()` in `src/external_ats.py`, applied in
|
| 18 |
+
`src/resume_v2_natural.py::_allocate_keywords` after `decide_includable_terms`.
|
| 19 |
+
Rejects: geography (word-boundary matched so `graphql` is safe from the `hq`
|
| 20 |
+
substring), the hiring company's own name (threaded in as `company`),
|
| 21 |
+
person names (cue-gated — a Title-Case bigram next to "Recruiter:" / "is the
|
| 22 |
+
CEO" / "— Manager", so job-title headings like "Forward Deployed" /
|
| 23 |
+
"Banking Integrations" are NOT false-flagged), corporate-entity nouns
|
| 24 |
+
(corporation, members, backbasers…), and job-board UI blurbs (actively
|
| 25 |
+
engaged, easy apply…).
|
| 26 |
+
- **V1 is deliberately untouched** — the shared `_is_term_like` /
|
| 27 |
+
`extract_external_keywords` path is unchanged, so V1 structured placement
|
| 28 |
+
behaves exactly as before. A regression test (`test_v1_extractor_unchanged`)
|
| 29 |
+
asserts `amsterdam` still appears in the V1 pool.
|
| 30 |
+
- **Result:** the Backbase JD's V2 pool now yields only genuine keywords
|
| 31 |
+
(`rest soap graphql webhooks`, `ipaas reconciliation agile`, `banking
|
| 32 |
+
integrations`, `gcp`) — zero names, geography, company, or boilerplate.
|
| 33 |
+
- **Scope:** extraction/filtering only. Honesty guardrails, injection grammar,
|
| 34 |
+
and fit-gate are NOT part of this phase (deferred).
|
| 35 |
+
- **Tests:** new `tests/test_jd_leak_filter.py` (6 tests); existing V1 (11) and
|
| 36 |
+
V2 (17) suites stay green — 34 total.
|
| 37 |
+
|
| 38 |
+
---
|
| 39 |
+
|
| 40 |
## 2026-06-25 — Designation + Filename Update
|
| 41 |
|
| 42 |
- **Designation**: Changed NxtWave role from "Internal Product Manager" to
|
README.md
CHANGED
|
@@ -424,6 +424,17 @@ genuine-sounding experience bullets. Reads as natural prose rather than keyword
|
|
| 424 |
lists; honesty boundaries preserved; slightly slower (one LLM call + compile).
|
| 425 |
Falls back to V1 comma placement if the LLM call fails.
|
| 426 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 427 |
### Selecting a version
|
| 428 |
|
| 429 |
| Surface | How to choose |
|
|
|
|
| 424 |
lists; honesty boundaries preserved; slightly slower (one LLM call + compile).
|
| 425 |
Falls back to V1 comma placement if the LLM call fails.
|
| 426 |
|
| 427 |
+
**Scraped-noise filter (V2-only).** Before allocation, V2 passes the keyword pool
|
| 428 |
+
through `filter_scraped_noise()` (`src/external_ats.py`), which drops tokens a raw
|
| 429 |
+
JD scrape drags in but that are never resume keywords: company geography
|
| 430 |
+
(Amsterdam, Latin America…), the hiring company's own name, executive/recruiter
|
| 431 |
+
and other person names (cue-gated, e.g. "Recruiter: …" / "… is the CEO"),
|
| 432 |
+
corporate-entity/boilerplate nouns (corporation, members, backbasers…), and
|
| 433 |
+
job-board page chrome ("actively engaged", "easy apply"…). Genuine skill/tool/
|
| 434 |
+
domain keywords (REST, SOAP, GraphQL, iPaaS, reconciliation, Agile, banking
|
| 435 |
+
integrations…) are preserved. **V1 placement is intentionally NOT filtered** —
|
| 436 |
+
this is scoped to V2 only.
|
| 437 |
+
|
| 438 |
### Selecting a version
|
| 439 |
|
| 440 |
| Surface | How to choose |
|
src/external_ats.py
CHANGED
|
@@ -81,6 +81,149 @@ _UI_SUBSTR = (
|
|
| 81 |
)
|
| 82 |
|
| 83 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 84 |
# Generic JD action-verbs / weak edge words trimmed from the START/END of a
|
| 85 |
# maximal run so an extracted phrase reads like a SKILL, not a sentence fragment
|
| 86 |
# ("lead end-to-end product lifecycle" → "end-to-end product lifecycle").
|
|
|
|
| 81 |
)
|
| 82 |
|
| 83 |
|
| 84 |
+
# ── V2-only scraped-noise filter (R27) ───────────────────────────────────────
|
| 85 |
+
# These three sets + helpers reject scraped, non-skill tokens (company
|
| 86 |
+
# geography, executive/recruiter names, role/boilerplate nouns) that the broad
|
| 87 |
+
# extractor cannot tell apart from real skills. They are applied ONLY in the V2
|
| 88 |
+
# allocation path via `filter_scraped_noise()`; V1 structured placement does NOT
|
| 89 |
+
# call them, so V1 behaviour is unchanged.
|
| 90 |
+
|
| 91 |
+
# Geography / location tokens that are never skills. Word-boundary matching is
|
| 92 |
+
# used for single-word entries so 'hq' never blocks 'graphql'; multi-word phrases
|
| 93 |
+
# ('latin america') match as substrings.
|
| 94 |
+
_GEO_EXTENDED: frozenset = frozenset({
|
| 95 |
+
"amsterdam", "netherlands", "europe", "africa", "asia", "asia-pacific",
|
| 96 |
+
"latin america", "latin", "north america", "south america", "middle east",
|
| 97 |
+
"apac", "emea", "headquartered", "headquarters", "hq", "worldwide",
|
| 98 |
+
"global", "offices", "office",
|
| 99 |
+
})
|
| 100 |
+
|
| 101 |
+
# Corporate-entity / boilerplate suffix words that travel with "About us" copy.
|
| 102 |
+
_ENTITY_NOISE: frozenset = frozenset({
|
| 103 |
+
"corporation", "incorporated", "subsidiary", "subsidiaries", "conglomerate",
|
| 104 |
+
})
|
| 105 |
+
|
| 106 |
+
# Job-board / LinkedIn page chrome that the DOM scrape drags in (the broad V1
|
| 107 |
+
# extractor's _UI_NOISE list doesn't cover these exact fragments).
|
| 108 |
+
_UI_BLURB: tuple = (
|
| 109 |
+
"actively engaged", "actively reviewing", "get alerts", "reach applicants",
|
| 110 |
+
"top applicant", "easy apply", "be an early applicant", "promoted by",
|
| 111 |
+
"set alert", "job alert", "people you can reach",
|
| 112 |
+
)
|
| 113 |
+
|
| 114 |
+
# Generic role/people nouns that indicate a person or HR context, not a skill.
|
| 115 |
+
_ROLE_NOISE: frozenset = frozenset({
|
| 116 |
+
"recruiter", "recruiters", "applicant", "applicants",
|
| 117 |
+
"founder", "co-founder", "cofounder", "ceo", "cto", "coo", "cfo",
|
| 118 |
+
"investor", "investors", "partner", "partners",
|
| 119 |
+
"member", "members", "shareholder", "shareholders",
|
| 120 |
+
"colleague", "colleagues", "peer", "peers", "backbasers",
|
| 121 |
+
"hire", "hiring", "hired",
|
| 122 |
+
})
|
| 123 |
+
|
| 124 |
+
# Common sentence/clause-starting words that appear Title Case but are NOT part
|
| 125 |
+
# of a person name — excluded from the two-word name pattern.
|
| 126 |
+
_SENTENCE_STARTERS: frozenset = frozenset({
|
| 127 |
+
"The", "A", "An", "We", "Our", "You", "This", "That", "These", "Those",
|
| 128 |
+
"For", "At", "In", "On", "By", "With", "To", "From", "And", "Or", "But",
|
| 129 |
+
"If", "As", "Is", "Are", "Was", "About", "During", "After", "Before",
|
| 130 |
+
"When", "Where", "Please", "Must", "Will", "Can", "Has", "Have", "Had",
|
| 131 |
+
"Be", "Do", "Does", "Did", "Should", "Would", "Could", "May", "Might",
|
| 132 |
+
"Not", "All", "Some", "Any", "Each", "Also", "Both", "More", "Most",
|
| 133 |
+
"How", "What", "Who", "Why", "Which",
|
| 134 |
+
})
|
| 135 |
+
|
| 136 |
+
# Cue words that appear immediately BEFORE a person name in a JD ("Recruiter:
|
| 137 |
+
# Mohammed Nayeem", "Founder Jouk Pleiter").
|
| 138 |
+
_NAME_CUE_BEFORE = (
|
| 139 |
+
r"recruiter|contact|hiring manager|posted by|submitted by|reach out to|"
|
| 140 |
+
r"founder|co-?founder|ceo|cto|coo|cfo|president|vice president|vp|"
|
| 141 |
+
r"director|head of [a-z ]+|lead"
|
| 142 |
+
)
|
| 143 |
+
|
| 144 |
+
|
| 145 |
+
def _contains_geo_token(term: str) -> bool:
|
| 146 |
+
"""True if `term` is or contains a geography/location token. Single-word geo
|
| 147 |
+
entries use word-boundary matching so the 'hq' substring never blocks
|
| 148 |
+
'graphql'; multi-word geo phrases (e.g. 'latin america') match as substrings."""
|
| 149 |
+
tl = (term or "").lower()
|
| 150 |
+
for g in _GEO_EXTENDED:
|
| 151 |
+
if g == tl:
|
| 152 |
+
return True
|
| 153 |
+
if " " in g and g in tl:
|
| 154 |
+
return True
|
| 155 |
+
if " " not in g and re.search(
|
| 156 |
+
r"(?<![a-z])" + re.escape(g) + r"(?![a-z])", tl
|
| 157 |
+
):
|
| 158 |
+
return True
|
| 159 |
+
return False
|
| 160 |
+
|
| 161 |
+
|
| 162 |
+
def _contains_person_name(term_lower: str, original_jd: str) -> bool:
|
| 163 |
+
"""True if `term_lower` contains a two-word Title-Case sequence that, in the
|
| 164 |
+
original JD, sits next to a person cue — a label before it ('Recruiter:') or a
|
| 165 |
+
predicate/appositive after it ('is the CEO', '— Account Manager', ', founder').
|
| 166 |
+
|
| 167 |
+
Context-gated on purpose: a bare Title-Case bigram is NOT enough (job-title and
|
| 168 |
+
domain headings like 'Forward Deployed' or 'Banking Integrations' are Title
|
| 169 |
+
Case too). Requiring a person cue avoids those false positives while still
|
| 170 |
+
catching real names embedded anywhere in a multi-word gram.
|
| 171 |
+
"""
|
| 172 |
+
words = term_lower.split()
|
| 173 |
+
if len(words) < 2:
|
| 174 |
+
return False
|
| 175 |
+
orig = original_jd or ""
|
| 176 |
+
for w1, w2 in zip(words, words[1:]):
|
| 177 |
+
cap1, cap2 = w1.capitalize(), w2.capitalize()
|
| 178 |
+
if cap1 in _SENTENCE_STARTERS or cap2 in _SENTENCE_STARTERS:
|
| 179 |
+
continue
|
| 180 |
+
name = re.escape(cap1) + r"\s+" + re.escape(cap2)
|
| 181 |
+
before = r"(?:" + _NAME_CUE_BEFORE + r")\s*[:\-–—]?\s+" + name
|
| 182 |
+
after = name + r"\s*(?:\bis\b|\bwas\b|,|[\-–—]|\bthe ceo\b|\bfounder\b)"
|
| 183 |
+
if re.search(before, orig, flags=re.I) or re.search(after, orig):
|
| 184 |
+
return True
|
| 185 |
+
return False
|
| 186 |
+
|
| 187 |
+
|
| 188 |
+
def _is_skill_like(term: str, original_jd: str = "", company: str = "") -> bool:
|
| 189 |
+
"""Return False (reject) if `term` is geography, a corporate-entity/role noun,
|
| 190 |
+
the company's own name, or a person name. This is the sole V2 rejection gate;
|
| 191 |
+
it does NOT touch the V1 extractor's stopword/UI/filler checks."""
|
| 192 |
+
tl = (term or "").strip().lower()
|
| 193 |
+
if not tl:
|
| 194 |
+
return True
|
| 195 |
+
if _is_ui_noise(tl):
|
| 196 |
+
return False
|
| 197 |
+
if any(b in tl for b in _UI_BLURB):
|
| 198 |
+
return False
|
| 199 |
+
if _contains_geo_token(tl):
|
| 200 |
+
return False
|
| 201 |
+
parts = tl.split()
|
| 202 |
+
if tl in _ROLE_NOISE or any(w in _ROLE_NOISE for w in parts):
|
| 203 |
+
return False
|
| 204 |
+
if any(w in _ENTITY_NOISE for w in parts):
|
| 205 |
+
return False
|
| 206 |
+
if company:
|
| 207 |
+
for ctok in re.findall(r"[a-z0-9]+", company.lower()):
|
| 208 |
+
if len(ctok) >= 4 and re.search(
|
| 209 |
+
r"\b" + re.escape(ctok) + r"(?:s|rs|ers)?\b", tl
|
| 210 |
+
):
|
| 211 |
+
return False
|
| 212 |
+
if original_jd and _contains_person_name(tl, original_jd):
|
| 213 |
+
return False
|
| 214 |
+
return True
|
| 215 |
+
|
| 216 |
+
|
| 217 |
+
def filter_scraped_noise(
|
| 218 |
+
terms: List[str], original_jd: str = "", company: str = ""
|
| 219 |
+
) -> List[str]:
|
| 220 |
+
"""V2-only: drop scraped non-skill tokens (geography, executive/recruiter
|
| 221 |
+
names, corporate-entity/boilerplate nouns, the company's own name) from an
|
| 222 |
+
already-extracted term list, preserving order. V1 does NOT call this — V1's
|
| 223 |
+
pool is unchanged by design."""
|
| 224 |
+
return [t for t in (terms or []) if _is_skill_like(t, original_jd, company)]
|
| 225 |
+
|
| 226 |
+
|
| 227 |
# Generic JD action-verbs / weak edge words trimmed from the START/END of a
|
| 228 |
# maximal run so an extracted phrase reads like a SKILL, not a sentence fragment
|
| 229 |
# ("lead end-to-end product lifecycle" → "end-to-end product lifecycle").
|
src/resume_v2_natural.py
CHANGED
|
@@ -30,7 +30,7 @@ from src.latex_resume import (
|
|
| 30 |
_is_hardcoded_resume,
|
| 31 |
_specialty_hit,
|
| 32 |
)
|
| 33 |
-
from src.external_ats import external_coverage
|
| 34 |
from src.candidate_fit import _REGULATED_CRED
|
| 35 |
|
| 36 |
log = logging.getLogger("resume_v2")
|
|
@@ -131,11 +131,17 @@ Return JSON:
|
|
| 131 |
|
| 132 |
# ── Keyword allocation (same waterfall as V1) ────────────────────────────────
|
| 133 |
|
| 134 |
-
def _allocate_keywords(jd_text: str, base_text: str, latex_src: str
|
|
|
|
| 135 |
"""Allocate JD keywords to sections using the exact V1 waterfall.
|
| 136 |
Returns (allocations_dict, decision_dict)."""
|
| 137 |
decision = decide_includable_terms(jd_text, base_text, maximum_ats_mode=True)
|
| 138 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 139 |
|
| 140 |
src_low = (latex_src or "").lower()
|
| 141 |
seen: set[str] = set()
|
|
@@ -409,7 +415,7 @@ def generate_v2(
|
|
| 409 |
base_text = latex_to_text(latex_src or "")
|
| 410 |
|
| 411 |
# 1. Allocate keywords (exact V1 waterfall)
|
| 412 |
-
allocations, decision = _allocate_keywords(jd_text, base_text, latex_src)
|
| 413 |
|
| 414 |
has_keywords = any(
|
| 415 |
(isinstance(v, list) and v) or (isinstance(v, dict) and any(vv for vv in v.values()))
|
|
|
|
| 30 |
_is_hardcoded_resume,
|
| 31 |
_specialty_hit,
|
| 32 |
)
|
| 33 |
+
from src.external_ats import external_coverage, filter_scraped_noise
|
| 34 |
from src.candidate_fit import _REGULATED_CRED
|
| 35 |
|
| 36 |
log = logging.getLogger("resume_v2")
|
|
|
|
| 131 |
|
| 132 |
# ── Keyword allocation (same waterfall as V1) ────────────────────────────────
|
| 133 |
|
| 134 |
+
def _allocate_keywords(jd_text: str, base_text: str, latex_src: str,
|
| 135 |
+
company: str = "") -> tuple[dict, dict]:
|
| 136 |
"""Allocate JD keywords to sections using the exact V1 waterfall.
|
| 137 |
Returns (allocations_dict, decision_dict)."""
|
| 138 |
decision = decide_includable_terms(jd_text, base_text, maximum_ats_mode=True)
|
| 139 |
+
# V2-only (R27): strip scraped non-skill noise — company geography, the
|
| 140 |
+
# company's own name, executive/recruiter names, role/boilerplate nouns — that
|
| 141 |
+
# the shared extractor leaves in the pool. V1 structured placement does NOT
|
| 142 |
+
# filter, so its behaviour is unchanged. `jd_text` is the original
|
| 143 |
+
# (pre-lowercase) JD, needed for the person-name cue heuristic.
|
| 144 |
+
includable = filter_scraped_noise(decision["includable"], jd_text, company)
|
| 145 |
|
| 146 |
src_low = (latex_src or "").lower()
|
| 147 |
seen: set[str] = set()
|
|
|
|
| 415 |
base_text = latex_to_text(latex_src or "")
|
| 416 |
|
| 417 |
# 1. Allocate keywords (exact V1 waterfall)
|
| 418 |
+
allocations, decision = _allocate_keywords(jd_text, base_text, latex_src, company)
|
| 419 |
|
| 420 |
has_keywords = any(
|
| 421 |
(isinstance(v, list) and v) or (isinstance(v, dict) and any(vv for vv in v.values()))
|
tests/test_jd_leak_filter.py
ADDED
|
@@ -0,0 +1,175 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""tests/test_jd_leak_filter.py — V2 JD-scraping-leak filter (R27).
|
| 2 |
+
|
| 3 |
+
Verifies the V2-ONLY scraped-noise filter: company geography, executive/recruiter
|
| 4 |
+
names, and role/boilerplate nouns are dropped from the V2 keyword pool, while
|
| 5 |
+
genuine skill/tool/domain keywords survive. Also asserts the change is scoped to
|
| 6 |
+
V2 — the shared V1 extractor (`extract_external_keywords`) is unchanged.
|
| 7 |
+
|
| 8 |
+
Run: python -m pytest tests/test_jd_leak_filter.py -x -q
|
| 9 |
+
"""
|
| 10 |
+
from __future__ import annotations
|
| 11 |
+
|
| 12 |
+
import os
|
| 13 |
+
import pytest
|
| 14 |
+
|
| 15 |
+
try:
|
| 16 |
+
from src.external_ats import (
|
| 17 |
+
filter_scraped_noise,
|
| 18 |
+
_is_skill_like,
|
| 19 |
+
extract_external_keywords,
|
| 20 |
+
)
|
| 21 |
+
except Exception as exc: # pragma: no cover - import guard
|
| 22 |
+
pytestmark = pytest.mark.skip(f"external_ats import unavailable: {exc}")
|
| 23 |
+
|
| 24 |
+
|
| 25 |
+
# Real Backbase-style JD: an "About us" block (geography, CEO, boilerplate) +
|
| 26 |
+
# a recruiter line, wrapped around the genuine role requirements.
|
| 27 |
+
BACKBASE_LIKE_JD = """
|
| 28 |
+
Forward Deployed Product Manager for Banking Integrations
|
| 29 |
+
|
| 30 |
+
About Backbase:
|
| 31 |
+
Backbase is headquartered in Amsterdam with offices in North America, Europe,
|
| 32 |
+
Middle East, Asia-Pacific, Africa, and Latin America. Jouk Pleiter is the CEO
|
| 33 |
+
and founder. Our worldwide network of backbasers, partners, peers, and members
|
| 34 |
+
share a bold vision.
|
| 35 |
+
|
| 36 |
+
About the Role:
|
| 37 |
+
We are looking for a Product Manager with experience in banking integrations,
|
| 38 |
+
API management, REST, SOAP, GraphQL, webhooks, iPaaS, reconciliation, Agile,
|
| 39 |
+
and stakeholder management. GCP experience preferred.
|
| 40 |
+
|
| 41 |
+
Recruiter: Mohammed Nayeem - Account Manager.
|
| 42 |
+
Actively engaged on LinkedIn Corporation to get alerts, reach applicants.
|
| 43 |
+
"""
|
| 44 |
+
|
| 45 |
+
# Tokens that MUST NOT survive the V2 filter.
|
| 46 |
+
BAD_TOKENS = {
|
| 47 |
+
"jouk pleiter",
|
| 48 |
+
"amsterdam",
|
| 49 |
+
"latin america",
|
| 50 |
+
"north america",
|
| 51 |
+
"headquartered",
|
| 52 |
+
"worldwide network",
|
| 53 |
+
"recruiter",
|
| 54 |
+
"members",
|
| 55 |
+
"mohammed nayeem",
|
| 56 |
+
"backbasers",
|
| 57 |
+
"corporation",
|
| 58 |
+
"linkedin corporation",
|
| 59 |
+
"actively engaged",
|
| 60 |
+
}
|
| 61 |
+
|
| 62 |
+
# The hiring company — its own name must never become a resume keyword.
|
| 63 |
+
COMPANY = "Backbase"
|
| 64 |
+
|
| 65 |
+
# Tokens that MUST survive (genuine skill/tool/domain keywords).
|
| 66 |
+
GOOD_TOKENS = {
|
| 67 |
+
"rest",
|
| 68 |
+
"soap",
|
| 69 |
+
"graphql",
|
| 70 |
+
"webhooks",
|
| 71 |
+
"reconciliation",
|
| 72 |
+
"agile",
|
| 73 |
+
"gcp",
|
| 74 |
+
"stakeholder management",
|
| 75 |
+
"banking integrations",
|
| 76 |
+
}
|
| 77 |
+
|
| 78 |
+
|
| 79 |
+
def _flatten_allocations(allocations: dict) -> set[str]:
|
| 80 |
+
out: set[str] = set()
|
| 81 |
+
for v in allocations.values():
|
| 82 |
+
if isinstance(v, list):
|
| 83 |
+
out.update(x.lower() for x in v)
|
| 84 |
+
elif isinstance(v, dict):
|
| 85 |
+
for pts in v.values():
|
| 86 |
+
out.update(x.lower() for x in pts)
|
| 87 |
+
return out
|
| 88 |
+
|
| 89 |
+
|
| 90 |
+
# ── filter_scraped_noise (the V2 gate, unit level) ───────────────────────────
|
| 91 |
+
|
| 92 |
+
def test_bad_tokens_absent():
|
| 93 |
+
"""Every known scraped token is rejected by the V2 filter."""
|
| 94 |
+
candidates = list(BAD_TOKENS) + list(GOOD_TOKENS)
|
| 95 |
+
survivors = {t.lower() for t in
|
| 96 |
+
filter_scraped_noise(candidates, BACKBASE_LIKE_JD, COMPANY)}
|
| 97 |
+
leaked = BAD_TOKENS & survivors
|
| 98 |
+
assert not leaked, f"scraped tokens leaked through V2 filter: {sorted(leaked)}"
|
| 99 |
+
|
| 100 |
+
|
| 101 |
+
def test_good_tokens_present():
|
| 102 |
+
"""No genuine keyword is over-filtered."""
|
| 103 |
+
candidates = list(BAD_TOKENS) + list(GOOD_TOKENS)
|
| 104 |
+
survivors = {t.lower() for t in
|
| 105 |
+
filter_scraped_noise(candidates, BACKBASE_LIKE_JD, COMPANY)}
|
| 106 |
+
dropped = GOOD_TOKENS - survivors
|
| 107 |
+
assert not dropped, f"genuine keywords wrongly dropped: {sorted(dropped)}"
|
| 108 |
+
|
| 109 |
+
|
| 110 |
+
def test_no_false_positives():
|
| 111 |
+
"""High-risk substrings/headings are kept (graphql vs 'hq'; Title-Case heading
|
| 112 |
+
'Banking Integrations' is a domain term, not a person name)."""
|
| 113 |
+
assert _is_skill_like("graphql", BACKBASE_LIKE_JD) # 'hq' must not match
|
| 114 |
+
assert _is_skill_like("ipaas", BACKBASE_LIKE_JD)
|
| 115 |
+
assert _is_skill_like("reconciliation", BACKBASE_LIKE_JD)
|
| 116 |
+
assert _is_skill_like("banking integrations", BACKBASE_LIKE_JD)
|
| 117 |
+
assert _is_skill_like("stakeholder management", BACKBASE_LIKE_JD)
|
| 118 |
+
# And the leak tokens are rejected at the unit level too.
|
| 119 |
+
assert not _is_skill_like("jouk pleiter", BACKBASE_LIKE_JD)
|
| 120 |
+
assert not _is_skill_like("amsterdam", BACKBASE_LIKE_JD)
|
| 121 |
+
assert not _is_skill_like("recruiter", BACKBASE_LIKE_JD)
|
| 122 |
+
assert not _is_skill_like("mohammed nayeem", BACKBASE_LIKE_JD)
|
| 123 |
+
|
| 124 |
+
|
| 125 |
+
# ── V2 allocation path (integration) ─────────────────────────────────────────
|
| 126 |
+
|
| 127 |
+
def test_v2_allocation_excludes_scraped():
|
| 128 |
+
"""End-to-end: V2's _allocate_keywords pool contains none of the bad tokens."""
|
| 129 |
+
try:
|
| 130 |
+
from src.resume_v2_natural import _allocate_keywords
|
| 131 |
+
from src.default_resume import get_default_resume_latex
|
| 132 |
+
from src.latex_resume import latex_to_text
|
| 133 |
+
except Exception as exc:
|
| 134 |
+
pytest.skip(f"V2 allocation deps unavailable: {exc}")
|
| 135 |
+
latex_src = get_default_resume_latex()
|
| 136 |
+
base_text = latex_to_text(latex_src)
|
| 137 |
+
allocations, _decision = _allocate_keywords(
|
| 138 |
+
BACKBASE_LIKE_JD, base_text, latex_src, COMPANY)
|
| 139 |
+
placed = _flatten_allocations(allocations)
|
| 140 |
+
leaked = {b for b in BAD_TOKENS if any(b == p or b in p for p in placed)}
|
| 141 |
+
assert not leaked, f"scraped tokens reached V2 allocations: {sorted(leaked)}"
|
| 142 |
+
|
| 143 |
+
|
| 144 |
+
# ── Scope guard: the change is V2-only ───────────────────────────────────────
|
| 145 |
+
|
| 146 |
+
def test_v1_extractor_unchanged():
|
| 147 |
+
"""The shared V1 extractor is NOT filtered — proving the fix is V2-scoped.
|
| 148 |
+
'amsterdam' still appears in extract_external_keywords output (V1 path)."""
|
| 149 |
+
pool = {t.lower() for t in extract_external_keywords(BACKBASE_LIKE_JD)}
|
| 150 |
+
assert "amsterdam" in pool, (
|
| 151 |
+
"V1 extractor output changed — filter must be V2-only, not in the shared path"
|
| 152 |
+
)
|
| 153 |
+
# But the V2 filter would remove it.
|
| 154 |
+
assert "amsterdam" not in {
|
| 155 |
+
t.lower() for t in filter_scraped_noise(list(pool), BACKBASE_LIKE_JD)
|
| 156 |
+
}
|
| 157 |
+
|
| 158 |
+
|
| 159 |
+
def test_real_jd_coverage_not_regressed():
|
| 160 |
+
"""A clean PM JD with no boilerplate keeps virtually all of its keyword pool
|
| 161 |
+
after V2 filtering (the filter only removes scraped non-skill noise)."""
|
| 162 |
+
clean_jd = (
|
| 163 |
+
"We are looking for a Product Manager. Responsibilities: roadmap planning. "
|
| 164 |
+
"Stakeholder management. A/B testing. Product analytics. Agile delivery. "
|
| 165 |
+
"Scrum ceremonies. User research. Feature prioritization. SQL dashboards. "
|
| 166 |
+
"OKRs and KPIs. PRD writing. Go-to-market strategy. Customer retention. "
|
| 167 |
+
"Funnel optimization. REST and GraphQL APIs. Data-driven decisions."
|
| 168 |
+
)
|
| 169 |
+
pool = extract_external_keywords(clean_jd)
|
| 170 |
+
filtered = filter_scraped_noise(pool, clean_jd, company="")
|
| 171 |
+
assert pool, "extractor produced no keywords for a clean JD"
|
| 172 |
+
# A clean JD has no geography/names/boilerplate, so the filter drops ~nothing.
|
| 173 |
+
assert len(filtered) >= int(0.9 * len(pool)), (
|
| 174 |
+
f"V2 filter over-pruned a clean JD: {len(filtered)}/{len(pool)} kept"
|
| 175 |
+
)
|