Spaces:
Sleeping
Sleeping
Commit Β·
fa74339
1
Parent(s): a52b643
feat(09-02): uncapped JD keyword extraction, honesty gate kept [R21]
Browse files- _is_term_like widened (drop suffix-artifact reject); no count cap
- extract_external_keywords gram capture edge-anchored (first+last
non-stop) + length 34->40, capturing more real phrases
- decide_includable_terms: in max_ats mode include every term except
the hard honesty gate (blocked + _specialty_hit); buzzwords allowed
in max mode per owner override. Non-max path unchanged.
- Honesty preserved: CISSP/PMP/CUDA/VP-of-Engineering still gated.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
.planning/phases/09-hardcoded-resume-keyword-placement/09-02-SUMMARY.md
ADDED
|
@@ -0,0 +1,19 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# 09-02 SUMMARY β Uncapped JD keyword extraction [R21]
|
| 2 |
+
|
| 3 |
+
Status: COMPLETE
|
| 4 |
+
|
| 5 |
+
## What changed
|
| 6 |
+
- `src/external_ats.py`
|
| 7 |
+
- `_is_term_like`: removed the suffix-artifact rejection (`endswith at/iz/ic`). Kept all genuine-noise guards (stop/filler/UI-noise/digit/len<3/>4-words). No count cap anywhere.
|
| 8 |
+
- `extract_external_keywords`: gram capture is now edge-anchored (first+last word non-stop/non-filler) instead of requiring ALL interior words non-stop; gram length cap 34β40. Captures more real phrases; `_is_term_like` inside `_add` still rejects all-stop grams + UI noise.
|
| 9 |
+
- `src/latex_resume.py`
|
| 10 |
+
- `decide_includable_terms`: in `maximum_ats_mode`, include EVERY missing term except the hard honesty gate (blocked + `_specialty_hit`). Buzzwords deliberately allowed in max mode (owner override, documented in code). Non-max path unchanged (buzzword + classify_fit gating preserved).
|
| 11 |
+
|
| 12 |
+
## Verification (all passed)
|
| 13 |
+
1. `_is_term_like`: accepts "automation"/"roadmap planning"; rejects "easy apply"/"the"/"innovation"
|
| 14 |
+
2. Extraction breadth: 22 terms on sample JD; UI noise ("easy apply","try premium") excluded; "roadmap" present
|
| 15 |
+
3. Honesty in max mode: CISSP/PMP/CUDA/VP-of-Engineering all gated; "roadmap" included
|
| 16 |
+
4. Non-max unchanged: "innovation" gated (buzzword), "roadmap" included
|
| 17 |
+
|
| 18 |
+
## Honesty boundary
|
| 19 |
+
NOT relaxed. blocked + `_specialty_hit` (deep-tech/engineering/senior/regulated-cert) reject in BOTH modes.
|
src/external_ats.py
CHANGED
|
@@ -101,7 +101,13 @@ def _is_ui_noise(t: str) -> bool:
|
|
| 101 |
|
| 102 |
|
| 103 |
def _is_term_like(t: str) -> bool:
|
| 104 |
-
"""A token/phrase that reads like a real skill/responsibility/domain term.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 105 |
t = t.strip()
|
| 106 |
if not t or t in _STOP or t in _FILLER:
|
| 107 |
return False
|
|
@@ -114,12 +120,10 @@ def _is_term_like(t: str) -> bool:
|
|
| 114 |
return False
|
| 115 |
if t.isdigit():
|
| 116 |
return False
|
| 117 |
-
# All words must be non-stop
|
| 118 |
for w in words:
|
| 119 |
if w in _STOP:
|
| 120 |
return False
|
| 121 |
-
if len(t) >= 5 and t.endswith(("at", "iz", "ic")) and len(words) == 1:
|
| 122 |
-
return False # "integrat", "automat", "operat"
|
| 123 |
return True
|
| 124 |
|
| 125 |
|
|
@@ -164,11 +168,15 @@ def extract_external_keywords(jd_text: str, extra: List[str] = None) -> List[str
|
|
| 164 |
for n in (3, 2):
|
| 165 |
for i in range(len(words) - n + 1):
|
| 166 |
gram = " ".join(words[i:i + n])
|
| 167 |
-
|
| 168 |
-
|
| 169 |
-
|
| 170 |
-
|
| 171 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 172 |
|
| 173 |
# 4. Pasted external terms (ground truth from a checker).
|
| 174 |
for t in (extra or []):
|
|
|
|
| 101 |
|
| 102 |
|
| 103 |
def _is_term_like(t: str) -> bool:
|
| 104 |
+
"""A token/phrase that reads like a real skill/responsibility/domain term.
|
| 105 |
+
|
| 106 |
+
Phase 9 (R21): intentionally WIDENED β "if in doubt, pick it up." We keep only
|
| 107 |
+
genuine-noise guards (stopwords, filler, UI/CTA overlay strings, pure digits,
|
| 108 |
+
too-short, >4 words). There is NO count cap; the honesty gate lives in
|
| 109 |
+
decide_includable_terms (specialty/cert/seniority/blocked), not here.
|
| 110 |
+
"""
|
| 111 |
t = t.strip()
|
| 112 |
if not t or t in _STOP or t in _FILLER:
|
| 113 |
return False
|
|
|
|
| 120 |
return False
|
| 121 |
if t.isdigit():
|
| 122 |
return False
|
| 123 |
+
# All words must be non-stop.
|
| 124 |
for w in words:
|
| 125 |
if w in _STOP:
|
| 126 |
return False
|
|
|
|
|
|
|
| 127 |
return True
|
| 128 |
|
| 129 |
|
|
|
|
| 168 |
for n in (3, 2):
|
| 169 |
for i in range(len(words) - n + 1):
|
| 170 |
gram = " ".join(words[i:i + n])
|
| 171 |
+
# R21: edge-anchored β keep a gram when its FIRST and LAST words are
|
| 172 |
+
# non-stop/non-filler (was: ALL interior words). Captures more real
|
| 173 |
+
# phrases (e.g. "go to market"); _add()'s _is_term_like still rejects
|
| 174 |
+
# all-stop grams and UI noise.
|
| 175 |
+
edge_ok = (words[i] not in _STOP and words[i] not in _FILLER
|
| 176 |
+
and words[i + n - 1] not in _STOP
|
| 177 |
+
and words[i + n - 1] not in _FILLER)
|
| 178 |
+
if edge_ok and jd_low.count(gram) >= 1 and len(gram) <= 40:
|
| 179 |
+
_add(gram)
|
| 180 |
|
| 181 |
# 4. Pasted external terms (ground truth from a checker).
|
| 182 |
for t in (extra or []):
|
src/latex_resume.py
CHANGED
|
@@ -176,19 +176,28 @@ def decide_includable_terms(
|
|
| 176 |
if t in blocked:
|
| 177 |
gated[t] = "blocked (user/vault)"
|
| 178 |
continue
|
| 179 |
-
if t in _BUZZWORDS:
|
| 180 |
-
gated[t] = "buzzword (checkers penalise)"
|
| 181 |
-
continue
|
| 182 |
# Phrase-level honesty guard: the broad external extractor yields
|
| 183 |
# multi-word grams ("cuda kernel programming") that the exact-match fit
|
| 184 |
# classifier misses. Block any phrase CONTAINING a specialized
|
| 185 |
-
# engineering / deep-tech / regulated-credential token.
|
| 186 |
# even in Maximum ATS Mode (unless the user explicitly confirmed it).
|
| 187 |
if t not in confirmed:
|
| 188 |
hit = _specialty_hit(t)
|
| 189 |
if hit:
|
| 190 |
gated[t] = f"block: {hit}"
|
| 191 |
continue
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 192 |
req = Requirement(term=t, category=_categorize(t))
|
| 193 |
verdict = classify_fit(
|
| 194 |
req, base_text, maximum_ats_mode=maximum_ats_mode, confirmed=confirmed,
|
|
|
|
| 176 |
if t in blocked:
|
| 177 |
gated[t] = "blocked (user/vault)"
|
| 178 |
continue
|
|
|
|
|
|
|
|
|
|
| 179 |
# Phrase-level honesty guard: the broad external extractor yields
|
| 180 |
# multi-word grams ("cuda kernel programming") that the exact-match fit
|
| 181 |
# classifier misses. Block any phrase CONTAINING a specialized
|
| 182 |
+
# engineering / deep-tech / regulated-credential token. NEVER relaxed,
|
| 183 |
# even in Maximum ATS Mode (unless the user explicitly confirmed it).
|
| 184 |
if t not in confirmed:
|
| 185 |
hit = _specialty_hit(t)
|
| 186 |
if hit:
|
| 187 |
gated[t] = f"block: {hit}"
|
| 188 |
continue
|
| 189 |
+
if maximum_ats_mode:
|
| 190 |
+
# R21 (Phase 9): uncapped β keep every plausible term. The honesty gate
|
| 191 |
+
# above (blocked + _specialty_hit/regulated/seniority) is the ONLY
|
| 192 |
+
# filter in max mode. Buzzwords are DELIBERATELY allowed here per the
|
| 193 |
+
# explicit owner override (see 09-CONTEXT.md "Explicit override");
|
| 194 |
+
# do NOT re-add the buzzword/classify_fit gate to this branch.
|
| 195 |
+
includable.append(t)
|
| 196 |
+
continue
|
| 197 |
+
# ββ Default (non-max) path: original conservative gating, unchanged. ββ
|
| 198 |
+
if t in _BUZZWORDS:
|
| 199 |
+
gated[t] = "buzzword (checkers penalise)"
|
| 200 |
+
continue
|
| 201 |
req = Requirement(term=t, category=_categorize(t))
|
| 202 |
verdict = classify_fit(
|
| 203 |
req, base_text, maximum_ats_mode=maximum_ats_mode, confirmed=confirmed,
|