saitejatirunagari Claude Opus 4.8 commited on
Commit
fa74339
Β·
1 Parent(s): a52b643

feat(09-02): uncapped JD keyword extraction, honesty gate kept [R21]

Browse files

- _is_term_like widened (drop suffix-artifact reject); no count cap
- extract_external_keywords gram capture edge-anchored (first+last
non-stop) + length 34->40, capturing more real phrases
- decide_includable_terms: in max_ats mode include every term except
the hard honesty gate (blocked + _specialty_hit); buzzwords allowed
in max mode per owner override. Non-max path unchanged.
- Honesty preserved: CISSP/PMP/CUDA/VP-of-Engineering still gated.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

.planning/phases/09-hardcoded-resume-keyword-placement/09-02-SUMMARY.md ADDED
@@ -0,0 +1,19 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # 09-02 SUMMARY β€” Uncapped JD keyword extraction [R21]
2
+
3
+ Status: COMPLETE
4
+
5
+ ## What changed
6
+ - `src/external_ats.py`
7
+ - `_is_term_like`: removed the suffix-artifact rejection (`endswith at/iz/ic`). Kept all genuine-noise guards (stop/filler/UI-noise/digit/len<3/>4-words). No count cap anywhere.
8
+ - `extract_external_keywords`: gram capture is now edge-anchored (first+last word non-stop/non-filler) instead of requiring ALL interior words non-stop; gram length cap 34β†’40. Captures more real phrases; `_is_term_like` inside `_add` still rejects all-stop grams + UI noise.
9
+ - `src/latex_resume.py`
10
+ - `decide_includable_terms`: in `maximum_ats_mode`, include EVERY missing term except the hard honesty gate (blocked + `_specialty_hit`). Buzzwords deliberately allowed in max mode (owner override, documented in code). Non-max path unchanged (buzzword + classify_fit gating preserved).
11
+
12
+ ## Verification (all passed)
13
+ 1. `_is_term_like`: accepts "automation"/"roadmap planning"; rejects "easy apply"/"the"/"innovation"
14
+ 2. Extraction breadth: 22 terms on sample JD; UI noise ("easy apply","try premium") excluded; "roadmap" present
15
+ 3. Honesty in max mode: CISSP/PMP/CUDA/VP-of-Engineering all gated; "roadmap" included
16
+ 4. Non-max unchanged: "innovation" gated (buzzword), "roadmap" included
17
+
18
+ ## Honesty boundary
19
+ NOT relaxed. blocked + `_specialty_hit` (deep-tech/engineering/senior/regulated-cert) reject in BOTH modes.
src/external_ats.py CHANGED
@@ -101,7 +101,13 @@ def _is_ui_noise(t: str) -> bool:
101
 
102
 
103
  def _is_term_like(t: str) -> bool:
104
- """A token/phrase that reads like a real skill/responsibility/domain term."""
 
 
 
 
 
 
105
  t = t.strip()
106
  if not t or t in _STOP or t in _FILLER:
107
  return False
@@ -114,12 +120,10 @@ def _is_term_like(t: str) -> bool:
114
  return False
115
  if t.isdigit():
116
  return False
117
- # All words must be non-stop and not lemmatizer artifacts.
118
  for w in words:
119
  if w in _STOP:
120
  return False
121
- if len(t) >= 5 and t.endswith(("at", "iz", "ic")) and len(words) == 1:
122
- return False # "integrat", "automat", "operat"
123
  return True
124
 
125
 
@@ -164,11 +168,15 @@ def extract_external_keywords(jd_text: str, extra: List[str] = None) -> List[str
164
  for n in (3, 2):
165
  for i in range(len(words) - n + 1):
166
  gram = " ".join(words[i:i + n])
167
- if all(w not in _STOP and w not in _FILLER for w in words[i:i + n]):
168
- # Only keep grams that recur or look like a skill phrase.
169
- if jd_low.count(gram) >= 1 and len(gram) <= 34:
170
- # Skip grams that are mostly filler-ish single words joined.
171
- _add(gram)
 
 
 
 
172
 
173
  # 4. Pasted external terms (ground truth from a checker).
174
  for t in (extra or []):
 
101
 
102
 
103
  def _is_term_like(t: str) -> bool:
104
+ """A token/phrase that reads like a real skill/responsibility/domain term.
105
+
106
+ Phase 9 (R21): intentionally WIDENED β€” "if in doubt, pick it up." We keep only
107
+ genuine-noise guards (stopwords, filler, UI/CTA overlay strings, pure digits,
108
+ too-short, >4 words). There is NO count cap; the honesty gate lives in
109
+ decide_includable_terms (specialty/cert/seniority/blocked), not here.
110
+ """
111
  t = t.strip()
112
  if not t or t in _STOP or t in _FILLER:
113
  return False
 
120
  return False
121
  if t.isdigit():
122
  return False
123
+ # All words must be non-stop.
124
  for w in words:
125
  if w in _STOP:
126
  return False
 
 
127
  return True
128
 
129
 
 
168
  for n in (3, 2):
169
  for i in range(len(words) - n + 1):
170
  gram = " ".join(words[i:i + n])
171
+ # R21: edge-anchored β€” keep a gram when its FIRST and LAST words are
172
+ # non-stop/non-filler (was: ALL interior words). Captures more real
173
+ # phrases (e.g. "go to market"); _add()'s _is_term_like still rejects
174
+ # all-stop grams and UI noise.
175
+ edge_ok = (words[i] not in _STOP and words[i] not in _FILLER
176
+ and words[i + n - 1] not in _STOP
177
+ and words[i + n - 1] not in _FILLER)
178
+ if edge_ok and jd_low.count(gram) >= 1 and len(gram) <= 40:
179
+ _add(gram)
180
 
181
  # 4. Pasted external terms (ground truth from a checker).
182
  for t in (extra or []):
src/latex_resume.py CHANGED
@@ -176,19 +176,28 @@ def decide_includable_terms(
176
  if t in blocked:
177
  gated[t] = "blocked (user/vault)"
178
  continue
179
- if t in _BUZZWORDS:
180
- gated[t] = "buzzword (checkers penalise)"
181
- continue
182
  # Phrase-level honesty guard: the broad external extractor yields
183
  # multi-word grams ("cuda kernel programming") that the exact-match fit
184
  # classifier misses. Block any phrase CONTAINING a specialized
185
- # engineering / deep-tech / regulated-credential token. Never relaxed,
186
  # even in Maximum ATS Mode (unless the user explicitly confirmed it).
187
  if t not in confirmed:
188
  hit = _specialty_hit(t)
189
  if hit:
190
  gated[t] = f"block: {hit}"
191
  continue
 
 
 
 
 
 
 
 
 
 
 
 
192
  req = Requirement(term=t, category=_categorize(t))
193
  verdict = classify_fit(
194
  req, base_text, maximum_ats_mode=maximum_ats_mode, confirmed=confirmed,
 
176
  if t in blocked:
177
  gated[t] = "blocked (user/vault)"
178
  continue
 
 
 
179
  # Phrase-level honesty guard: the broad external extractor yields
180
  # multi-word grams ("cuda kernel programming") that the exact-match fit
181
  # classifier misses. Block any phrase CONTAINING a specialized
182
+ # engineering / deep-tech / regulated-credential token. NEVER relaxed,
183
  # even in Maximum ATS Mode (unless the user explicitly confirmed it).
184
  if t not in confirmed:
185
  hit = _specialty_hit(t)
186
  if hit:
187
  gated[t] = f"block: {hit}"
188
  continue
189
+ if maximum_ats_mode:
190
+ # R21 (Phase 9): uncapped β€” keep every plausible term. The honesty gate
191
+ # above (blocked + _specialty_hit/regulated/seniority) is the ONLY
192
+ # filter in max mode. Buzzwords are DELIBERATELY allowed here per the
193
+ # explicit owner override (see 09-CONTEXT.md "Explicit override");
194
+ # do NOT re-add the buzzword/classify_fit gate to this branch.
195
+ includable.append(t)
196
+ continue
197
+ # ── Default (non-max) path: original conservative gating, unchanged. ──
198
+ if t in _BUZZWORDS:
199
+ gated[t] = "buzzword (checkers penalise)"
200
+ continue
201
  req = Requirement(term=t, category=_categorize(t))
202
  verdict = classify_fit(
203
  req, base_text, maximum_ats_mode=maximum_ats_mode, confirmed=confirmed,