saitejatirunagari Claude Opus 4.6 commited on
Commit
1ec12ef
·
1 Parent(s): 9badb7a

feat(phase-11): kill JD-scraping leak in keyword pool — V2-only

Browse files

Scoped to V2 per request (plan assumed shared V1+V2 chokepoint). New
filter_scraped_noise()/_is_skill_like() in src/external_ats.py, applied only in
src/resume_v2_natural.py::_allocate_keywords. Rejects geography (graphql safe from
'hq'), the hiring company's own name, cue-gated person names (so job-title
headings survive), corporate-entity/role nouns, and job-board UI blurbs. V1
extractor path untouched. New tests/test_jd_leak_filter.py (6 tests); V1 (11) + V2
(17) suites green — 34 total.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

.planning/REQUIREMENTS.md CHANGED
@@ -106,3 +106,6 @@ V2 must NOT append keyword lists. Instead, fan out across the existing model poo
106
 
107
  ## R26: Label V1 + documentation + default-safe
108
  Explicitly label the current behavior as "V1" in README.md and HISTORY.md, document V2, and ensure V1 stays the default so existing flows are unaffected if no version is specified.
 
 
 
 
106
 
107
  ## R26: Label V1 + documentation + default-safe
108
  Explicitly label the current behavior as "V1" in README.md and HISTORY.md, document V2, and ensure V1 stays the default so existing flows are unaffected if no version is specified.
109
+
110
+ ## R27: Kill the JD-scraping leak in the keyword extractor
111
+ The keyword extraction layer (the code that turns a raw JD into the keyword pool consumed by BOTH V1 structured placement and V2 sentence generation) is leaking scraped, non-skill text directly into the candidate's resume. Real evidence from a generated Backbase resume: the Skills "Other" block contained the company CEO's name (`jouk pleiter`), geography boilerplate (`headquartered, amsterdam, north america europe, middle east asia-pacific africa, latin america`), About-us marketing copy (`backbase helped financial institutions … bold climate action … worldwide network, backbasers partners`), and recruiter-blurb fragments (`re interested, likely, recruiter, don, show`); the injected summary contained scraped LinkedIn UI text (`actively engaged on LinkedIn Corporation to get alerts, reach applicants, and stay a top applicant`) and a stray collaborator name (`Mohammed Nayeem`). The extractor MUST NOT emit: company marketing/boilerplate copy, person/named-entity tokens (executives, recruiters, collaborators), location/geography lists, multi-word sentence fragments, or generic recruiter/UI-blurb tokens. Only genuine, role-relevant skill / tool / domain keywords survive. Scope is the extraction + filtering layer ONLY — NOT the honesty guardrails, injection grammar, or fit-gate (deferred). Success: a JD containing company boilerplate + a CEO name + a geography list yields a keyword pool with zero of those tokens, and a regenerated resume's Skills "Other" line and summary contain no scraped non-skill text.
.planning/phases/11-jd-scraping-leak-fix/11-01-SUMMARY.md ADDED
@@ -0,0 +1,61 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ phase: 11-jd-scraping-leak-fix
3
+ plan: 01
4
+ type: summary
5
+ status: DONE
6
+ requirements: [R27]
7
+ deviation: "Scoped to V2-only per user instruction (plan assumed shared V1+V2 chokepoint)"
8
+ ---
9
+
10
+ # 11-01 Summary: Kill the JD-Scraping Leak (V2-only)
11
+
12
+ ## Deviation from plan
13
+ The verified plan fixed the **shared** chokepoint (`_is_term_like` in
14
+ `external_ats.py`), which would have changed both V1 and V2. The user instructed
15
+ "ensure this only goes in V2." Implementation was re-scoped: the filter is a
16
+ standalone function applied **only in the V2 allocation path**; the shared V1
17
+ extractor is untouched.
18
+
19
+ ## What was built
20
+
21
+ ### src/external_ats.py (additive — definitions only, not wired into V1)
22
+ - `_GEO_EXTENDED`, `_ENTITY_NOISE`, `_ROLE_NOISE`, `_SENTENCE_STARTERS`,
23
+ `_UI_BLURB`, `_NAME_CUE_BEFORE`
24
+ - `_contains_geo_token()` — word-boundary single-word match (so `hq` never
25
+ blocks `graphql`) + substring match for multi-word geo phrases
26
+ - `_contains_person_name()` — **cue-gated**: a Title-Case bigram is a name only
27
+ if it sits next to a person cue ("Recruiter: …", "… is the CEO", "— Manager").
28
+ Scans all adjacent pairs, so a name embedded in a longer gram is caught. Avoids
29
+ false-flagging job-title/domain headings ("Forward Deployed", "Banking
30
+ Integrations", "Product Manager")
31
+ - `_is_skill_like(term, original_jd, company)` — the rejection gate: UI noise,
32
+ UI blurbs, geography, role/entity nouns, the hiring company's own name
33
+ (word-boundary + s/rs/ers suffix), and cue-gated person names
34
+ - `filter_scraped_noise(terms, original_jd, company)` — public V2 entry point
35
+
36
+ ### src/resume_v2_natural.py (V2 path only)
37
+ - `_allocate_keywords()` gained a `company` param; calls
38
+ `filter_scraped_noise(decision["includable"], jd_text, company)` right after
39
+ `decide_includable_terms`
40
+ - `generate_v2()` passes `company` into `_allocate_keywords`
41
+
42
+ ### tests/test_jd_leak_filter.py (new, 6 tests)
43
+ - `test_bad_tokens_absent`, `test_good_tokens_present`, `test_no_false_positives`
44
+ - `test_v2_allocation_excludes_scraped` — end-to-end via `_allocate_keywords`
45
+ - `test_v1_extractor_unchanged` — **proves V2-only**: `amsterdam` still in the V1
46
+ `extract_external_keywords` pool, removed only by `filter_scraped_noise`
47
+ - `test_real_jd_coverage_not_regressed` — clean JD keeps ≥90% of its pool
48
+
49
+ ### Docs
50
+ - README "Generation Modes" → V2 scraped-noise filter subsection (notes V1 not filtered)
51
+ - HISTORY → Phase 11 entry
52
+
53
+ ## Verification
54
+ - `pytest tests/test_jd_leak_filter.py tests/test_structured_placement.py tests/test_resume_v2.py -q` → **34 passed**
55
+ - Real Backbase JD, V2 pool after fix: `['banking integrations api management',
56
+ 'forward deployed product manager', 'gcp', 'ipaas reconciliation agile',
57
+ 'rest soap graphql webhooks']` — zero names/geography/company/boilerplate.
58
+
59
+ ## Out of scope (deferred, unchanged)
60
+ Honesty guardrails (`_REGULATED_CRED`/`_specialty_hit`), injection grammar/
61
+ coherence, fit-gate, upstream scraper. V1 placement behavior.
.planning/phases/11-jd-scraping-leak-fix/11-CONTEXT.md ADDED
@@ -0,0 +1,78 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ phase: 11-jd-scraping-leak-fix
3
+ type: context
4
+ captured: 2026-06-25
5
+ source: real generated resume (Backbase Forward Deployed PM) + user-confirmed scope
6
+ ---
7
+
8
+ # Phase 11 Context — Kill the JD-Scraping Leak
9
+
10
+ ## Origin
11
+ A live resume generated for a **Backbase "Forward Deployed Product Manager for
12
+ Banking Integrations"** role (LinkedIn job 4415276270) was analyzed as if by the
13
+ hiring HR head. It would be rejected instantly — not for tuning reasons but
14
+ because the **keyword extractor leaked raw scraped JD/company text directly into
15
+ the resume**, producing fabrication-looking content and obvious automated-stuffing
16
+ signals.
17
+
18
+ ## The actual leaked tokens (verbatim evidence)
19
+ From `Saiteja_Tirunagari_Backbase_Resume.tex`:
20
+
21
+ 1. **Skills "Other" block** (`% ats-skills-other`, line 266) contained:
22
+ - Executive name: `jouk pleiter` (Backbase's CEO)
23
+ - Geography list: `headquartered, amsterdam, north america europe, middle east asia-pacific africa, latin america`
24
+ - About-us marketing copy: `backbase helped financial institutions, institutions stand, re taking bold climate, action helping future-proof, planet alongside, worldwide network, backbasers partners, peers, members, share`
25
+ - Recruiter/UI-blurb fragments: `re interested, likely, recruiter, don, show`
26
+ - Other sentence fragments: `things grand central, voice, client inside`
27
+
28
+ 2. **Injected summary** (`% ats-injected-start`, line 148) contained:
29
+ - Scraped LinkedIn UI text: `actively engaged on LinkedIn Corporation to get alerts, reach applicants, and stay a top applicant in competitive PM roles`
30
+ - Unsupported claim tokens: `GCP, banking integrations`
31
+
32
+ 3. **Injected bullets** referenced a stray collaborator name: `Mohammed Nayeem`
33
+ (bleed-through from unrelated profile/JD data).
34
+
35
+ ## Root-cause hypothesis (to confirm in research)
36
+ The JD→keyword extraction layer is treating the ENTIRE scraped page (company
37
+ "About us", CEO bio, recruiter blurb, LinkedIn chrome) as JD body, then tokenizing
38
+ it into the keyword pool with no named-entity / boilerplate / fragment filtering.
39
+ This pool feeds BOTH:
40
+ - V1 structured placement (`src/latex_resume.py` — `decide_includable_terms`,
41
+ inject into `% ats-skills-other` + `% ats-injected-start`)
42
+ - V2 sentence generation (`src/resume_v2_natural.py` — `_allocate_keywords`)
43
+
44
+ ## Scope decision (USER-CONFIRMED)
45
+ **IN scope:** the extraction + filtering layer only — stop non-skill tokens from
46
+ ever entering the keyword pool.
47
+
48
+ The extractor MUST NOT emit:
49
+ - company marketing / boilerplate copy
50
+ - person / named-entity tokens (executives, recruiters, collaborators)
51
+ - location / geography lists
52
+ - multi-word sentence fragments
53
+ - generic recruiter / UI-blurb tokens
54
+
55
+ Only genuine, role-relevant **skill / tool / domain** keywords survive.
56
+
57
+ **OUT of scope (explicitly deferred by the user — do NOT touch):**
58
+ - honesty guardrails (`_REGULATED_CRED`, `_specialty_hit`)
59
+ - grammar/coherence of the injected sentences
60
+ - fit-gate / domain-mismatch down-ranking
61
+ - the upstream JD scraper itself (we filter at extraction, not at scrape)
62
+
63
+ ## Constraints
64
+ - Fix at the SHARED extraction layer so both V1 and V2 benefit from one change.
65
+ - Do NOT over-filter: legitimate keywords (REST, SOAP, GraphQL, iPaaS, Agile,
66
+ banking-integration domain terms) must still survive — no regression in real
67
+ keyword coverage.
68
+ - Existing tests must stay green: `tests/test_structured_placement.py` (V1, 11
69
+ tests), `tests/test_resume_v2.py` (V2, 17 tests).
70
+ - Project rule [[feedback_docs_update]]: update README.md + HISTORY.md.
71
+ - Honesty principle [[ats_no_stuffing_principle]]: this is the mechanism behind
72
+ "keyword dumps get stripped" — fixing extraction is the upstream cure.
73
+
74
+ ## Success (acceptance)
75
+ A JD containing company boilerplate + a CEO name + a geography list yields a
76
+ keyword pool with **zero** of those tokens; a regenerated resume's Skills "Other"
77
+ line and summary contain no scraped non-skill text; true-keyword coverage does
78
+ not regress; both V1 and V2 test suites pass.
HISTORY.md CHANGED
@@ -4,6 +4,39 @@ A running log of everything built, fixed, and changed. Most recent first.
4
 
5
  ---
6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7
  ## 2026-06-25 — Designation + Filename Update
8
 
9
  - **Designation**: Changed NxtWave role from "Internal Product Manager" to
 
4
 
5
  ---
6
 
7
+ ## 2026-06-25 — Phase 11: Kill the JD-Scraping Leak (V2-only)
8
+
9
+ A live Backbase resume exposed the keyword extractor leaking scraped, non-skill
10
+ text into the resume: the Skills "Other" block held the CEO's name (`jouk
11
+ pleiter`), geography (`amsterdam`, `latin america`), About-us copy (`worldwide
12
+ network`, `backbasers partners`), and recruiter fragments; the summary held
13
+ LinkedIn UI chrome (`actively engaged on LinkedIn Corporation…`) and a stray
14
+ collaborator name (`mohammed nayeem`).
15
+
16
+ - **Fix (scoped to V2 only, per request):** new `filter_scraped_noise()` +
17
+ `_is_skill_like()` in `src/external_ats.py`, applied in
18
+ `src/resume_v2_natural.py::_allocate_keywords` after `decide_includable_terms`.
19
+ Rejects: geography (word-boundary matched so `graphql` is safe from the `hq`
20
+ substring), the hiring company's own name (threaded in as `company`),
21
+ person names (cue-gated — a Title-Case bigram next to "Recruiter:" / "is the
22
+ CEO" / "— Manager", so job-title headings like "Forward Deployed" /
23
+ "Banking Integrations" are NOT false-flagged), corporate-entity nouns
24
+ (corporation, members, backbasers…), and job-board UI blurbs (actively
25
+ engaged, easy apply…).
26
+ - **V1 is deliberately untouched** — the shared `_is_term_like` /
27
+ `extract_external_keywords` path is unchanged, so V1 structured placement
28
+ behaves exactly as before. A regression test (`test_v1_extractor_unchanged`)
29
+ asserts `amsterdam` still appears in the V1 pool.
30
+ - **Result:** the Backbase JD's V2 pool now yields only genuine keywords
31
+ (`rest soap graphql webhooks`, `ipaas reconciliation agile`, `banking
32
+ integrations`, `gcp`) — zero names, geography, company, or boilerplate.
33
+ - **Scope:** extraction/filtering only. Honesty guardrails, injection grammar,
34
+ and fit-gate are NOT part of this phase (deferred).
35
+ - **Tests:** new `tests/test_jd_leak_filter.py` (6 tests); existing V1 (11) and
36
+ V2 (17) suites stay green — 34 total.
37
+
38
+ ---
39
+
40
  ## 2026-06-25 — Designation + Filename Update
41
 
42
  - **Designation**: Changed NxtWave role from "Internal Product Manager" to
README.md CHANGED
@@ -424,6 +424,17 @@ genuine-sounding experience bullets. Reads as natural prose rather than keyword
424
  lists; honesty boundaries preserved; slightly slower (one LLM call + compile).
425
  Falls back to V1 comma placement if the LLM call fails.
426
 
 
 
 
 
 
 
 
 
 
 
 
427
  ### Selecting a version
428
 
429
  | Surface | How to choose |
 
424
  lists; honesty boundaries preserved; slightly slower (one LLM call + compile).
425
  Falls back to V1 comma placement if the LLM call fails.
426
 
427
+ **Scraped-noise filter (V2-only).** Before allocation, V2 passes the keyword pool
428
+ through `filter_scraped_noise()` (`src/external_ats.py`), which drops tokens a raw
429
+ JD scrape drags in but that are never resume keywords: company geography
430
+ (Amsterdam, Latin America…), the hiring company's own name, executive/recruiter
431
+ and other person names (cue-gated, e.g. "Recruiter: …" / "… is the CEO"),
432
+ corporate-entity/boilerplate nouns (corporation, members, backbasers…), and
433
+ job-board page chrome ("actively engaged", "easy apply"…). Genuine skill/tool/
434
+ domain keywords (REST, SOAP, GraphQL, iPaaS, reconciliation, Agile, banking
435
+ integrations…) are preserved. **V1 placement is intentionally NOT filtered** —
436
+ this is scoped to V2 only.
437
+
438
  ### Selecting a version
439
 
440
  | Surface | How to choose |
src/external_ats.py CHANGED
@@ -81,6 +81,149 @@ _UI_SUBSTR = (
81
  )
82
 
83
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
84
  # Generic JD action-verbs / weak edge words trimmed from the START/END of a
85
  # maximal run so an extracted phrase reads like a SKILL, not a sentence fragment
86
  # ("lead end-to-end product lifecycle" → "end-to-end product lifecycle").
 
81
  )
82
 
83
 
84
+ # ── V2-only scraped-noise filter (R27) ───────────────────────────────────────
85
+ # These three sets + helpers reject scraped, non-skill tokens (company
86
+ # geography, executive/recruiter names, role/boilerplate nouns) that the broad
87
+ # extractor cannot tell apart from real skills. They are applied ONLY in the V2
88
+ # allocation path via `filter_scraped_noise()`; V1 structured placement does NOT
89
+ # call them, so V1 behaviour is unchanged.
90
+
91
+ # Geography / location tokens that are never skills. Word-boundary matching is
92
+ # used for single-word entries so 'hq' never blocks 'graphql'; multi-word phrases
93
+ # ('latin america') match as substrings.
94
+ _GEO_EXTENDED: frozenset = frozenset({
95
+ "amsterdam", "netherlands", "europe", "africa", "asia", "asia-pacific",
96
+ "latin america", "latin", "north america", "south america", "middle east",
97
+ "apac", "emea", "headquartered", "headquarters", "hq", "worldwide",
98
+ "global", "offices", "office",
99
+ })
100
+
101
+ # Corporate-entity / boilerplate suffix words that travel with "About us" copy.
102
+ _ENTITY_NOISE: frozenset = frozenset({
103
+ "corporation", "incorporated", "subsidiary", "subsidiaries", "conglomerate",
104
+ })
105
+
106
+ # Job-board / LinkedIn page chrome that the DOM scrape drags in (the broad V1
107
+ # extractor's _UI_NOISE list doesn't cover these exact fragments).
108
+ _UI_BLURB: tuple = (
109
+ "actively engaged", "actively reviewing", "get alerts", "reach applicants",
110
+ "top applicant", "easy apply", "be an early applicant", "promoted by",
111
+ "set alert", "job alert", "people you can reach",
112
+ )
113
+
114
+ # Generic role/people nouns that indicate a person or HR context, not a skill.
115
+ _ROLE_NOISE: frozenset = frozenset({
116
+ "recruiter", "recruiters", "applicant", "applicants",
117
+ "founder", "co-founder", "cofounder", "ceo", "cto", "coo", "cfo",
118
+ "investor", "investors", "partner", "partners",
119
+ "member", "members", "shareholder", "shareholders",
120
+ "colleague", "colleagues", "peer", "peers", "backbasers",
121
+ "hire", "hiring", "hired",
122
+ })
123
+
124
+ # Common sentence/clause-starting words that appear Title Case but are NOT part
125
+ # of a person name — excluded from the two-word name pattern.
126
+ _SENTENCE_STARTERS: frozenset = frozenset({
127
+ "The", "A", "An", "We", "Our", "You", "This", "That", "These", "Those",
128
+ "For", "At", "In", "On", "By", "With", "To", "From", "And", "Or", "But",
129
+ "If", "As", "Is", "Are", "Was", "About", "During", "After", "Before",
130
+ "When", "Where", "Please", "Must", "Will", "Can", "Has", "Have", "Had",
131
+ "Be", "Do", "Does", "Did", "Should", "Would", "Could", "May", "Might",
132
+ "Not", "All", "Some", "Any", "Each", "Also", "Both", "More", "Most",
133
+ "How", "What", "Who", "Why", "Which",
134
+ })
135
+
136
+ # Cue words that appear immediately BEFORE a person name in a JD ("Recruiter:
137
+ # Mohammed Nayeem", "Founder Jouk Pleiter").
138
+ _NAME_CUE_BEFORE = (
139
+ r"recruiter|contact|hiring manager|posted by|submitted by|reach out to|"
140
+ r"founder|co-?founder|ceo|cto|coo|cfo|president|vice president|vp|"
141
+ r"director|head of [a-z ]+|lead"
142
+ )
143
+
144
+
145
+ def _contains_geo_token(term: str) -> bool:
146
+ """True if `term` is or contains a geography/location token. Single-word geo
147
+ entries use word-boundary matching so the 'hq' substring never blocks
148
+ 'graphql'; multi-word geo phrases (e.g. 'latin america') match as substrings."""
149
+ tl = (term or "").lower()
150
+ for g in _GEO_EXTENDED:
151
+ if g == tl:
152
+ return True
153
+ if " " in g and g in tl:
154
+ return True
155
+ if " " not in g and re.search(
156
+ r"(?<![a-z])" + re.escape(g) + r"(?![a-z])", tl
157
+ ):
158
+ return True
159
+ return False
160
+
161
+
162
+ def _contains_person_name(term_lower: str, original_jd: str) -> bool:
163
+ """True if `term_lower` contains a two-word Title-Case sequence that, in the
164
+ original JD, sits next to a person cue — a label before it ('Recruiter:') or a
165
+ predicate/appositive after it ('is the CEO', '— Account Manager', ', founder').
166
+
167
+ Context-gated on purpose: a bare Title-Case bigram is NOT enough (job-title and
168
+ domain headings like 'Forward Deployed' or 'Banking Integrations' are Title
169
+ Case too). Requiring a person cue avoids those false positives while still
170
+ catching real names embedded anywhere in a multi-word gram.
171
+ """
172
+ words = term_lower.split()
173
+ if len(words) < 2:
174
+ return False
175
+ orig = original_jd or ""
176
+ for w1, w2 in zip(words, words[1:]):
177
+ cap1, cap2 = w1.capitalize(), w2.capitalize()
178
+ if cap1 in _SENTENCE_STARTERS or cap2 in _SENTENCE_STARTERS:
179
+ continue
180
+ name = re.escape(cap1) + r"\s+" + re.escape(cap2)
181
+ before = r"(?:" + _NAME_CUE_BEFORE + r")\s*[:\-–—]?\s+" + name
182
+ after = name + r"\s*(?:\bis\b|\bwas\b|,|[\-–—]|\bthe ceo\b|\bfounder\b)"
183
+ if re.search(before, orig, flags=re.I) or re.search(after, orig):
184
+ return True
185
+ return False
186
+
187
+
188
+ def _is_skill_like(term: str, original_jd: str = "", company: str = "") -> bool:
189
+ """Return False (reject) if `term` is geography, a corporate-entity/role noun,
190
+ the company's own name, or a person name. This is the sole V2 rejection gate;
191
+ it does NOT touch the V1 extractor's stopword/UI/filler checks."""
192
+ tl = (term or "").strip().lower()
193
+ if not tl:
194
+ return True
195
+ if _is_ui_noise(tl):
196
+ return False
197
+ if any(b in tl for b in _UI_BLURB):
198
+ return False
199
+ if _contains_geo_token(tl):
200
+ return False
201
+ parts = tl.split()
202
+ if tl in _ROLE_NOISE or any(w in _ROLE_NOISE for w in parts):
203
+ return False
204
+ if any(w in _ENTITY_NOISE for w in parts):
205
+ return False
206
+ if company:
207
+ for ctok in re.findall(r"[a-z0-9]+", company.lower()):
208
+ if len(ctok) >= 4 and re.search(
209
+ r"\b" + re.escape(ctok) + r"(?:s|rs|ers)?\b", tl
210
+ ):
211
+ return False
212
+ if original_jd and _contains_person_name(tl, original_jd):
213
+ return False
214
+ return True
215
+
216
+
217
+ def filter_scraped_noise(
218
+ terms: List[str], original_jd: str = "", company: str = ""
219
+ ) -> List[str]:
220
+ """V2-only: drop scraped non-skill tokens (geography, executive/recruiter
221
+ names, corporate-entity/boilerplate nouns, the company's own name) from an
222
+ already-extracted term list, preserving order. V1 does NOT call this — V1's
223
+ pool is unchanged by design."""
224
+ return [t for t in (terms or []) if _is_skill_like(t, original_jd, company)]
225
+
226
+
227
  # Generic JD action-verbs / weak edge words trimmed from the START/END of a
228
  # maximal run so an extracted phrase reads like a SKILL, not a sentence fragment
229
  # ("lead end-to-end product lifecycle" → "end-to-end product lifecycle").
src/resume_v2_natural.py CHANGED
@@ -30,7 +30,7 @@ from src.latex_resume import (
30
  _is_hardcoded_resume,
31
  _specialty_hit,
32
  )
33
- from src.external_ats import external_coverage
34
  from src.candidate_fit import _REGULATED_CRED
35
 
36
  log = logging.getLogger("resume_v2")
@@ -131,11 +131,17 @@ Return JSON:
131
 
132
  # ── Keyword allocation (same waterfall as V1) ────────────────────────────────
133
 
134
- def _allocate_keywords(jd_text: str, base_text: str, latex_src: str) -> tuple[dict, dict]:
 
135
  """Allocate JD keywords to sections using the exact V1 waterfall.
136
  Returns (allocations_dict, decision_dict)."""
137
  decision = decide_includable_terms(jd_text, base_text, maximum_ats_mode=True)
138
- includable = decision["includable"]
 
 
 
 
 
139
 
140
  src_low = (latex_src or "").lower()
141
  seen: set[str] = set()
@@ -409,7 +415,7 @@ def generate_v2(
409
  base_text = latex_to_text(latex_src or "")
410
 
411
  # 1. Allocate keywords (exact V1 waterfall)
412
- allocations, decision = _allocate_keywords(jd_text, base_text, latex_src)
413
 
414
  has_keywords = any(
415
  (isinstance(v, list) and v) or (isinstance(v, dict) and any(vv for vv in v.values()))
 
30
  _is_hardcoded_resume,
31
  _specialty_hit,
32
  )
33
+ from src.external_ats import external_coverage, filter_scraped_noise
34
  from src.candidate_fit import _REGULATED_CRED
35
 
36
  log = logging.getLogger("resume_v2")
 
131
 
132
  # ── Keyword allocation (same waterfall as V1) ────────────────────────────────
133
 
134
+ def _allocate_keywords(jd_text: str, base_text: str, latex_src: str,
135
+ company: str = "") -> tuple[dict, dict]:
136
  """Allocate JD keywords to sections using the exact V1 waterfall.
137
  Returns (allocations_dict, decision_dict)."""
138
  decision = decide_includable_terms(jd_text, base_text, maximum_ats_mode=True)
139
+ # V2-only (R27): strip scraped non-skill noise — company geography, the
140
+ # company's own name, executive/recruiter names, role/boilerplate nouns — that
141
+ # the shared extractor leaves in the pool. V1 structured placement does NOT
142
+ # filter, so its behaviour is unchanged. `jd_text` is the original
143
+ # (pre-lowercase) JD, needed for the person-name cue heuristic.
144
+ includable = filter_scraped_noise(decision["includable"], jd_text, company)
145
 
146
  src_low = (latex_src or "").lower()
147
  seen: set[str] = set()
 
415
  base_text = latex_to_text(latex_src or "")
416
 
417
  # 1. Allocate keywords (exact V1 waterfall)
418
+ allocations, decision = _allocate_keywords(jd_text, base_text, latex_src, company)
419
 
420
  has_keywords = any(
421
  (isinstance(v, list) and v) or (isinstance(v, dict) and any(vv for vv in v.values()))
tests/test_jd_leak_filter.py ADDED
@@ -0,0 +1,175 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """tests/test_jd_leak_filter.py — V2 JD-scraping-leak filter (R27).
2
+
3
+ Verifies the V2-ONLY scraped-noise filter: company geography, executive/recruiter
4
+ names, and role/boilerplate nouns are dropped from the V2 keyword pool, while
5
+ genuine skill/tool/domain keywords survive. Also asserts the change is scoped to
6
+ V2 — the shared V1 extractor (`extract_external_keywords`) is unchanged.
7
+
8
+ Run: python -m pytest tests/test_jd_leak_filter.py -x -q
9
+ """
10
+ from __future__ import annotations
11
+
12
+ import os
13
+ import pytest
14
+
15
+ try:
16
+ from src.external_ats import (
17
+ filter_scraped_noise,
18
+ _is_skill_like,
19
+ extract_external_keywords,
20
+ )
21
+ except Exception as exc: # pragma: no cover - import guard
22
+ pytestmark = pytest.mark.skip(f"external_ats import unavailable: {exc}")
23
+
24
+
25
+ # Real Backbase-style JD: an "About us" block (geography, CEO, boilerplate) +
26
+ # a recruiter line, wrapped around the genuine role requirements.
27
+ BACKBASE_LIKE_JD = """
28
+ Forward Deployed Product Manager for Banking Integrations
29
+
30
+ About Backbase:
31
+ Backbase is headquartered in Amsterdam with offices in North America, Europe,
32
+ Middle East, Asia-Pacific, Africa, and Latin America. Jouk Pleiter is the CEO
33
+ and founder. Our worldwide network of backbasers, partners, peers, and members
34
+ share a bold vision.
35
+
36
+ About the Role:
37
+ We are looking for a Product Manager with experience in banking integrations,
38
+ API management, REST, SOAP, GraphQL, webhooks, iPaaS, reconciliation, Agile,
39
+ and stakeholder management. GCP experience preferred.
40
+
41
+ Recruiter: Mohammed Nayeem - Account Manager.
42
+ Actively engaged on LinkedIn Corporation to get alerts, reach applicants.
43
+ """
44
+
45
+ # Tokens that MUST NOT survive the V2 filter.
46
+ BAD_TOKENS = {
47
+ "jouk pleiter",
48
+ "amsterdam",
49
+ "latin america",
50
+ "north america",
51
+ "headquartered",
52
+ "worldwide network",
53
+ "recruiter",
54
+ "members",
55
+ "mohammed nayeem",
56
+ "backbasers",
57
+ "corporation",
58
+ "linkedin corporation",
59
+ "actively engaged",
60
+ }
61
+
62
+ # The hiring company — its own name must never become a resume keyword.
63
+ COMPANY = "Backbase"
64
+
65
+ # Tokens that MUST survive (genuine skill/tool/domain keywords).
66
+ GOOD_TOKENS = {
67
+ "rest",
68
+ "soap",
69
+ "graphql",
70
+ "webhooks",
71
+ "reconciliation",
72
+ "agile",
73
+ "gcp",
74
+ "stakeholder management",
75
+ "banking integrations",
76
+ }
77
+
78
+
79
+ def _flatten_allocations(allocations: dict) -> set[str]:
80
+ out: set[str] = set()
81
+ for v in allocations.values():
82
+ if isinstance(v, list):
83
+ out.update(x.lower() for x in v)
84
+ elif isinstance(v, dict):
85
+ for pts in v.values():
86
+ out.update(x.lower() for x in pts)
87
+ return out
88
+
89
+
90
+ # ── filter_scraped_noise (the V2 gate, unit level) ───────────────────────────
91
+
92
+ def test_bad_tokens_absent():
93
+ """Every known scraped token is rejected by the V2 filter."""
94
+ candidates = list(BAD_TOKENS) + list(GOOD_TOKENS)
95
+ survivors = {t.lower() for t in
96
+ filter_scraped_noise(candidates, BACKBASE_LIKE_JD, COMPANY)}
97
+ leaked = BAD_TOKENS & survivors
98
+ assert not leaked, f"scraped tokens leaked through V2 filter: {sorted(leaked)}"
99
+
100
+
101
+ def test_good_tokens_present():
102
+ """No genuine keyword is over-filtered."""
103
+ candidates = list(BAD_TOKENS) + list(GOOD_TOKENS)
104
+ survivors = {t.lower() for t in
105
+ filter_scraped_noise(candidates, BACKBASE_LIKE_JD, COMPANY)}
106
+ dropped = GOOD_TOKENS - survivors
107
+ assert not dropped, f"genuine keywords wrongly dropped: {sorted(dropped)}"
108
+
109
+
110
+ def test_no_false_positives():
111
+ """High-risk substrings/headings are kept (graphql vs 'hq'; Title-Case heading
112
+ 'Banking Integrations' is a domain term, not a person name)."""
113
+ assert _is_skill_like("graphql", BACKBASE_LIKE_JD) # 'hq' must not match
114
+ assert _is_skill_like("ipaas", BACKBASE_LIKE_JD)
115
+ assert _is_skill_like("reconciliation", BACKBASE_LIKE_JD)
116
+ assert _is_skill_like("banking integrations", BACKBASE_LIKE_JD)
117
+ assert _is_skill_like("stakeholder management", BACKBASE_LIKE_JD)
118
+ # And the leak tokens are rejected at the unit level too.
119
+ assert not _is_skill_like("jouk pleiter", BACKBASE_LIKE_JD)
120
+ assert not _is_skill_like("amsterdam", BACKBASE_LIKE_JD)
121
+ assert not _is_skill_like("recruiter", BACKBASE_LIKE_JD)
122
+ assert not _is_skill_like("mohammed nayeem", BACKBASE_LIKE_JD)
123
+
124
+
125
+ # ── V2 allocation path (integration) ─────────────────────────────────────────
126
+
127
+ def test_v2_allocation_excludes_scraped():
128
+ """End-to-end: V2's _allocate_keywords pool contains none of the bad tokens."""
129
+ try:
130
+ from src.resume_v2_natural import _allocate_keywords
131
+ from src.default_resume import get_default_resume_latex
132
+ from src.latex_resume import latex_to_text
133
+ except Exception as exc:
134
+ pytest.skip(f"V2 allocation deps unavailable: {exc}")
135
+ latex_src = get_default_resume_latex()
136
+ base_text = latex_to_text(latex_src)
137
+ allocations, _decision = _allocate_keywords(
138
+ BACKBASE_LIKE_JD, base_text, latex_src, COMPANY)
139
+ placed = _flatten_allocations(allocations)
140
+ leaked = {b for b in BAD_TOKENS if any(b == p or b in p for p in placed)}
141
+ assert not leaked, f"scraped tokens reached V2 allocations: {sorted(leaked)}"
142
+
143
+
144
+ # ── Scope guard: the change is V2-only ───────────────────────────────────────
145
+
146
+ def test_v1_extractor_unchanged():
147
+ """The shared V1 extractor is NOT filtered — proving the fix is V2-scoped.
148
+ 'amsterdam' still appears in extract_external_keywords output (V1 path)."""
149
+ pool = {t.lower() for t in extract_external_keywords(BACKBASE_LIKE_JD)}
150
+ assert "amsterdam" in pool, (
151
+ "V1 extractor output changed — filter must be V2-only, not in the shared path"
152
+ )
153
+ # But the V2 filter would remove it.
154
+ assert "amsterdam" not in {
155
+ t.lower() for t in filter_scraped_noise(list(pool), BACKBASE_LIKE_JD)
156
+ }
157
+
158
+
159
+ def test_real_jd_coverage_not_regressed():
160
+ """A clean PM JD with no boilerplate keeps virtually all of its keyword pool
161
+ after V2 filtering (the filter only removes scraped non-skill noise)."""
162
+ clean_jd = (
163
+ "We are looking for a Product Manager. Responsibilities: roadmap planning. "
164
+ "Stakeholder management. A/B testing. Product analytics. Agile delivery. "
165
+ "Scrum ceremonies. User research. Feature prioritization. SQL dashboards. "
166
+ "OKRs and KPIs. PRD writing. Go-to-market strategy. Customer retention. "
167
+ "Funnel optimization. REST and GraphQL APIs. Data-driven decisions."
168
+ )
169
+ pool = extract_external_keywords(clean_jd)
170
+ filtered = filter_scraped_noise(pool, clean_jd, company="")
171
+ assert pool, "extractor produced no keywords for a clean JD"
172
+ # A clean JD has no geography/names/boilerplate, so the filter drops ~nothing.
173
+ assert len(filtered) >= int(0.9 * len(pool)), (
174
+ f"V2 filter over-pruned a clean JD: {len(filtered)}/{len(pool)} kept"
175
+ )