JAA-ATS-Tool / HISTORY.md
saitejatirunagari's picture
feat(ui): merge legacy platform checkboxes into unified platform selector
53c490d
|
Raw
History Blame
25.5 kB
# Project History β€” Job Automation Agent
A running log of everything built, fixed, and changed. Most recent first.
---
## 2026-06-13 β€” Unified Platform Selector + ATS + HTML Rendering Fixes
### Changes
- **Unified platform selector**: Merged the 6 legacy checkboxes ("🌐 Job Platforms") and the grouped ever-jobs selector ("🌐 ever-jobs Platforms") into a single "🌐 Job Platforms" section. One place to search all 170 platforms. Selecting LinkedIn/Indeed/Glassdoor/Remotive/WeWorkRemotely/Naukri still routes to their dedicated high-quality scrapers; everything else goes through EverJobsScraper.
- **ATS min_score default**: Changed slider default from 6 to 1 β€” LLM resumes now generated for ALL jobs regardless of score.
- **HTML rendering fix**: Switched all 5 `st.markdown(..., unsafe_allow_html=True)` calls to `st.html()` β€” fixes raw `<span>`/`<a>` tags showing as plain text in job cards (Streamlit 1.45+ regression).
### Modified Files
- `ui.py` β€” removed 6 legacy checkboxes, renamed section label, updated platforms_cfg, updated pipeline routing to use unified `all_platforms` key
---
## 2026-06-13 β€” Phase 1: ever-jobs Integration (160+ Platforms)
### New Features
- **160+ job platforms** via ever-jobs REST API integration (was 5 platforms)
- **Grouped platform selector** in UI: Search Boards / ATS Platforms / Company Pages with st.multiselect search
- **India-focused defaults**: 10 platforms pre-selected (LinkedIn, Naukri, Indeed, Glassdoor, Google, BDJobs, Internshala, Bayt, IIMJobs, Foundit)
- **Content fingerprint dedup**: SHA-256 of (title+company) catches cross-platform duplicates where same job appears on LinkedIn AND Greenhouse with different URLs
- **Performance warning**: UI shows warning when >30 platforms selected
### New Files
- `src/ever_jobs_bridge/__init__.py` β€” package init
- `src/ever_jobs_bridge/server.py` β€” Docker/npm server lifecycle (start/stop/health)
- `src/ever_jobs_bridge/client.py` β€” HTTP client for POST /api/jobs/search
- `src/ever_jobs_bridge/mapper.py` β€” IJob JSON β†’ Job dataclass field mapper
- `src/ever_jobs_bridge/platforms.py` β€” 170 platform catalog with group metadata
- `src/scrapers/ever_jobs.py` β€” EverJobsScraper extending BaseScraper
- `vendor/ever-jobs/` β€” ever-jobs NestJS monorepo (cloned, gitignored)
### Modified Files
- `src/job_history.py` β€” added content_fp column + is_duplicate_by_content() function
- `config.py` β€” added EVER_JOBS config block
- `ui.py` β€” grouped platform selector + EverJobsScraper pipeline wiring + ever_jobs step
- `requirements.txt` β€” added rapidfuzz>=3.0
- `.gitignore` β€” added vendor/
### R3 ATS Finding (Definitive)
ever-jobs "ATS" = Applicant Tracking System platforms that companies use to POST jobs
(Greenhouse, Lever, Workday). This is NOT resume scoring.
Our `src/ats_scorer.py` (70% JD keyword match + 30% resume quality) is the correct
resume ATS scoring system and is UNCHANGED. No modifications to ats_scorer.py are needed.
### Backward Compatibility
All existing scrapers (LinkedIn, Indeed, Glassdoor, Remotive, WeWorkRemotely) are UNTOUCHED.
Pipeline flow is unchanged β€” ever-jobs is an additive parallel path.
---
## Session 10 β€” 2026-06-13
### New: 2 additional job platforms (Remotive + We Work Remotely)
- **`src/scrapers/remotive.py`** β€” Remotive.io public JSON API. No auth needed.
Fetches WFH/remote PM jobs globally (India-eligible: "Worldwide" / APAC filter).
- **`src/scrapers/weworkremotely.py`** β€” We Work Remotely RSS feed scraper.
Free-to-scrape, good volume of remote PM roles.
- Both expose `get_details_bulk()` (no-op, descriptions come with the listing).
- Both appear as checkboxes in the new UI; step-skip if unchecked.
### Fixed: max_resumes slider removed β€” all jobs now get a resume
Previously `max_resumes` slider (default 15) silently capped LLM resumes even
when 30–40 jobs were fetched. Fixed by passing `max_llm_resumes=len(assessed_jobs)`
(effectively no cap). Every eligible job now gets an LLM-tailored resume.
### Fixed: platform cap is now total-per-platform, not per-query
Old code applied `max_results=N` per roleΓ—location query. With 3 roles Γ— 3 locations
you could get 9 Γ— 15 = 135 from one platform β€” far more than the user intended.
New code: the outer loop breaks once `platform_jobs` reaches `max_jobs_per_platform`,
and the per-query `max_results` is set to `remaining = cap - len(platform_jobs)`.
### Fixed: Google Sheets error messages are now informative
- `FileNotFoundError` (no credentials) now emits a clear "run setup_google.py" hint
- Full error text (up to 120 chars) logged to the live UI log, not just the file log
- A "Google Sheet status" indicator (βœ“/⚠) shown in the Configure section before run
### New: run history (save + load past runs)
- **`src/run_history.py`** β€” saves each completed run as JSON in
`data/output/run_history/run_YYYY-MM-DD_HH-MM-SS.json`.
Summary fields stored without jobs for fast listing; full jobs on load.
- History is auto-saved at the end of every pipeline run.
- UI "Load" button restores any past run's results to the active session without
rerunning the pipeline.
### New: complete UI redesign (ui.py)
- **No sidebar** β€” all controls now live inline in the main area.
- **History panel** β€” top-right "πŸ“œ History" button opens a panel listing all
past runs with stats (jobs, high-priority count, ATS before/after). Click "Load"
to restore any run.
- **Configure section** β€” expandable card with resume upload, roles, locations,
platform checkboxes, days, max-per-platform, and min score. Google Sheet
status shown inline.
- **Start button** β€” centered, prominent, full-width.
- **Step timeline** β€” CSS grid layout (auto-fill columns), fits all platforms.
- **Results tab β€” job cards** β€” top 10 shown as visual cards (title, company,
ATS before/after, salary, apply link). Switch to "Full Table" for all jobs.
- **Download fix** β€” zip now contains only the current run's date subfolder (not
all historical date folders). Eliminates the "90 files for 30 jobs" confusion
(per run: 30 DOCX + 30 PDF = 60 files as expected).
- **Metrics row** β€” Total | High | Medium | LLM Resumes | PDFs | Avg ATS After.
- Welcome state shown when no results are loaded yet.
### Fixed: test_mode β†’ False in config.py
Was accidentally left `True`, capping the pipeline at 10 jobs per test run.
---
## Session 9 β€” 2026-06-13
### Fixed: UI stuck at "0% β€” Starting…" while pipeline ran fine in background
**Symptom:** Click Start β†’ UI shows 0% and all steps "Waiting…" forever, but the
console/logs show the pipeline scraping, assessing 41 jobs, and generating
resumes at 91–94% ATS. Users clicked Start again thinking it was dead β†’ duplicate
pipeline threads (Thread-8 + Thread-17 in the logs).
**Root cause:** `_progress_q = queue.Queue()` was created at MODULE level in
ui.py with a comment claiming module globals survive reruns. They do NOT β€”
Streamlit re-executes the entry script top-to-bottom on EVERY rerun, creating a
brand-new empty Queue each time. The background thread kept writing progress to
the original queue; the UI drain loop polled the new empty one. Nothing ever
arrived.
**Fix (ui.py):**
- Queue now lives in `st.session_state["progress_q"]` β€” the only store that
survives reruns within a session
- `run_pipeline` receives the queue as an explicit default arg (`_q=_progress_q`)
and shadows the module helpers, so the thread always writes to the queue the
drain loop reads β€” even across reruns and multiple sessions
- `st.session_state["current_log_file"]` was being set FROM the background
thread (the "missing ScriptRunContext" warning, silently broken) β€” now sent
through the queue as a `("logfile", path)` message handled by the drain loop
**Verified with Streamlit AppTest:** queue identity preserved across reruns;
clicked Start in the test harness β€” UI received 7 log messages, step cards
updated (resume βœ… β†’ profile βœ… β†’ linkedin ⏳), progress bar at 15%.
**Files changed:** `ui.py`, `HISTORY.md`
---
## Session 8 β€” 2026-06-12
### Major performance + quality overhaul: parallel resumes, PDF output, full JD fetching
**Root causes of "taking lot of time, not going forward":**
1. LLM resumes generated ONE at a time (50–150s each Γ— 30 = up to an hour, UI frozen)
2. Indeed launched a full Chromium browser PER job description (~10s overhead each)
3. Glassdoor NEVER fetched descriptions (no detail method existed)
4. LinkedIn `job_id` regex broken β€” LinkedIn switched to slug URLs
(`/jobs/view/title-at-company-4423634421`), so ALL detail fetches 404'd β†’ no JDs
5. UI capped search to 3 roles Γ— 2 locations
**Fixes:**
- `src/resume_customizer.py` β€” LLM resumes now generated IN PARALLEL via
ThreadPoolExecutor (6 workers, round-robin across phase2 model API keys).
Per-resume `progress_cb` streams live status to the UI.
- `src/scrapers/linkedin.py` β€” fixed job_id extraction (slug URLs); new
`get_details_bulk()` fetches ALL descriptions with 4 parallel HTTP workers
- `src/scrapers/indeed.py` β€” new `get_details_bulk()`: ONE browser session for
all job descriptions instead of one browser per job
- `src/scrapers/glassdoor.py` β€” new `get_details_bulk()` with Cloudflare-challenge
wait + JSON-LD JobPosting parsing (Glassdoor still intermittent β€” bot-hostile)
- `ui.py` β€” searches ALL selected roles Γ— locations (caps removed); cross-platform
dedup by (title, company) in addition to URL; live per-resume progress
**ATS quality fixes (tailored resumes were sometimes scoring LOWER than original):**
- `src/llm_client.py` β€” validates LLM customization (summary >50 chars, β‰₯5 skills),
retries once, unwraps JSON arrays, max_tokens 3000β†’4000
- `resume_customizer.py` β€” optimization loop now: scores with same extra_kw as
final report Β· skips empty customizations Β· retries fall back to Kimi Β· rewrites
BEST attempt to disk (was keeping last) Β· GUARANTEE: if LLM result scores below
the original resume, ships keyword-injected template instead (After β‰₯ Before always)
- `_inject_missing_keywords()` rewritten β€” now injects the ACTUAL missing JD
keywords (was injecting generic PM keywords that didn't move the JD-match score)
**PDF output (new):**
- `src/pdf_writer.py` — DOCX→PDF: one Word COM session per batch on Windows
(perfect fidelity), reportlab re-render fallback on Linux/HF Spaces
- Every resume now saved as both `.docx` and `.pdf` in `data/output/resumes/YYYY-MM-DD/`
- UI: PDF + DOCX download buttons per job; zip download includes PDFs
- `requirements.txt`: + reportlab, docx2pdf (win32 only)
**Files changed:** `src/pdf_writer.py` (new), `src/resume_customizer.py`,
`src/llm_client.py`, `src/scrapers/linkedin.py`, `src/scrapers/indeed.py`,
`src/scrapers/glassdoor.py`, `ui.py`, `requirements.txt`, `README.md`, `HISTORY.md`
---
## Session 7 β€” 2026-06-12
### File-based logging system + Logs tab in UI
**Problem:** Pipeline was failing on HF Spaces with no way to see why. Queue-based live log only showed last 30 messages and swallowed full tracebacks.
**What was built:**
**`src/app_logger.py`** β€” New centralized logger:
- Writes every run to `data/logs/run_YYYY-MM-DD_HH-MM-SS.log`
- Captures ALL Python logging output (INFO, WARNING, ERROR, DEBUG)
- Redirects stdout/stderr via `_TeeStream` so `print()` and Playwright output are also captured
- In-memory ring buffer (500 lines) for UI access without file I/O
- `list_log_files()` returns all previous runs, newest first
**`ui.py`** changes:
- New **πŸ“‹ Logs** tab (5th tab)
- Color-coded viewer: errors=red, warnings=yellow, INFO done=green, info=blue
- Slider to show 50–500 lines
- Toggle to show/hide DEBUG lines
- Auto-refresh every 2s while pipeline is running
- Download button for raw `.log` file
- Previous run selector to load any past log
- Error/warning counts in footer
- Pipeline thread now calls `app_logger.setup()` at start β†’ creates timestamped log file
- Every scrape attempt logged with role + location + raw result count
- Full tracebacks on scrape errors (`logging.error(..., traceback)`)
- Fatal pipeline exceptions logged in full, not truncated to 400 chars
- `current_log_file` added to session state defaults
**`Dockerfile`** β€” Added `data/logs` to `mkdir -p` list
**Files changed:** `src/app_logger.py` (new), `ui.py`, `Dockerfile`, `HISTORY.md`, `README.md`
---
## Session 6 β€” 2026-06-11
### GitHub push + Hugging Face Spaces deployment prep
**Code pushed to GitHub:** https://github.com/saitejatiru/JAA-ATS-Tool
**HF Spaces files added:**
- `README.md` β€” prepended YAML frontmatter (`sdk: streamlit`, `app_file: ui.py`)
- `packages.txt` β€” Chromium system dependencies for Playwright on Linux
- `.gitignore` β€” excludes secrets (`google_token.json`, `.env`, resumes, output data)
- `.env.example` β€” documents all 9 NVIDIA API keys + Google Sheet ID
- `requirements.txt` β€” added `gspread`, `google-auth`, `google-auth-oauthlib`, `google-api-python-client`
**`ui.py` changes for HF Spaces:**
- Playwright install: `@st.cache_resource` function installs Chromium once per server lifetime
- Google credentials bootstrap: reads `GOOGLE_CREDENTIALS_JSON` env var and writes to `google_credentials.json` on startup
**Files changed:** `README.md`, `requirements.txt`, `packages.txt`, `.gitignore`, `.env.example`, `ui.py`
---
## Session 5 β€” 2026-06-11
### ATS Before/After in Excel + Verbose Resume Error Logging
**Excel reporter fixed:**
- Added `ATS Before (%)`, `ATS After (%)`, `ATS Improvement` columns to all sheets (was completely missing)
- Column order: Relevance Score β†’ ATS Before β†’ ATS After β†’ ATS Improvement β†’ Skills Match β†’ …
- `_pct()` helper: shows `"45%"` or `"β€”"` for null; improvement shows `"+37pp"` or `"β€”"`
- Column indices for score badge (9), URL hyperlink (23), priority color (15) updated to match new order
**Resume error visibility:**
- Added explicit `tqdm.write()` on success: `"βœ“ LLM resume: Google β†’ ATS 45% β†’ 82% (+37pp)"`
- Added `traceback.format_exc()` on failure so exact error is visible in the terminal
- Fallback ATS scoring (original resume score) always runs on failure so sheet never shows blank
**Confirmed working (run completed 2026-06-11 11:16):**
- 7 LLM-tailored + 2 template resumes generated in `data/output/resumes/2026-06-11/`
- Google Sheet updated with all 10 jobs
- Files: Google_Product Manager I Ads.docx, Instagram, Workday, Giga, Denave, Tessera, Latinem
**Files changed:** `src/excel_reporter.py`, `src/resume_customizer.py`
---
## Session 4 β€” 2026-06-11
### ATS Before/After Fix + Best Resume Prompt
**ATS Before/After not showing β€” root causes fixed:**
1. `score_resume()` was calling Kimi AGAIN (via `fast_model_cfg`) during ATS scoring β€” after already using Kimi for 9 resume generations, rate limits caused silent failures and blank scores. Fixed: removed `fast_model_cfg` from scoring calls; use pre-extracted keywords from assessment phase only.
2. On resume generation failure, `ats_score_before/after` was never set at all. Fixed: fallback block now always computes and stores ATS scores even if DOCX generation fails.
**Best ATS resume β€” prompt redesigned:**
- Old prompt: generic instructions, 1500 char JD limit, 2000 token output
- New prompt:
- Explicit mandatory keyword list with instruction "MUST include ALL of these"
- Rules enforce: exact JD language mirroring, action verbs on every bullet, quantified metrics required
- JD limit raised to 2000 chars, resume to 2500 chars
- Output tokens raised to 3000 (room for full detailed resume)
- 15 core competencies (was 12)
- More specific bullet format: "β€’ Led X resulting in Y% improvement"
**Profile extraction speed fix:**
- Step 2 was blocked on GLM 5.1 (~234s). Now tries Kimi-K2.6 (~5s) first via `extract_profile_summary_fast(cfg, ...)` with fallback to GLM.
- Added `LLMClient.extract_profile_summary_fast(cfg, resume_text)` method.
**Files changed:** `src/llm_client.py`, `src/resume_customizer.py`, `main.py`
---
## Session 3 β€” 2026-06-11
### Streamlit UI Fixes + LLM Resume Root-Cause Fix
**4 issues addressed:**
| Issue | Fix |
|-------|-----|
| LLM resumes = 0 | Root cause: `ATSScorer` class imported but never existed β†’ silent `ImportError`. Fixed by replacing with `score_resume()` function. Also fixed `PM_DOMAIN_KEYWORDS` β†’ `PM_BASE_KEYWORDS + PM_TOOLS` |
| Fast model for resume generation | Added `LLMClient._call_with_cfg()` + `customize_resume_fast(cfg, ...)`. Now uses Kimi-K2.6 (~5s) instead of GLM (~234s) |
| Date-based local resume folders | Resumes now save to `data/output/resumes/YYYY-MM-DD/`. No more Google Drive upload |
| Sheet headers missing | `gsheets.py` now detects missing header row and inserts at row 1 using `ws.insert_row()` even when data already exists |
| Test limit | 5 β†’ 10 jobs |
**Streamlit UI updated:**
- Fixed `customize_for_jobs()` parameter mismatch (`min_score` β†’ `min_score_for_llm`, `max_count` β†’ `max_llm_resumes`)
- Resume zip download now scans all date subfolders (`Path.rglob("*.docx")`)
- Results table now shows **ATS Before, ATS After, ATS Gain** columns
- Job Details tab shows ATS before/after inline
- `fast_model_cfg` wired into UI pipeline (Kimi-K2.6 for LLM keywords + resume tailoring)
**To launch UI:**
```powershell
streamlit run ui.py
# Opens at http://localhost:8501
```
---
## Session 2 β€” 2026-06-11
### Test Run Completed Successfully βœ…
**Results:**
- LinkedIn 60 + Indeed 18 + Glassdoor 13 jobs scraped (capped to 5 in test mode)
- Assessment: **16 seconds** for 5 jobs (Kimi K2.6, single batch)
- Top job: Associate Product Manager (Adtech) at MakeMyTrip β€” Score 8/10
- Google Sheet updated: https://docs.google.com/spreadsheets/d/1Ehxt3eortehbtySdtgcSrMhCqmxIMUAmvRqSkII0HJk/edit
- Excel saved: `data/output/reports/job_report.xlsx`
- 5 jobs marked in dedup store (SQLite) β€” won't reappear next run
**Bugs found during test run:**
1. `bulk_mark_seen` AttributeError β€” `Job` dataclass doesn't have `.get()`. Fixed with `isinstance(job, dict)` + `getattr()`.
2. Drive upload: `'Client' object has no attribute 'auth'` β€” gspread doesn't expose Drive API directly. **Still pending fix.**
3. LLM resumes = 0 β€” resume customization calling GLM (234s), timing out silently. **Still pending fix** (need to switch to Kimi/Step).
---
### ATS Scoring β€” Rebuilt from Scratch
**Problem:** Original ATS scored resume quality (structural), not job-description match. A generic resume scored the same for any job.
**Solution:** Resume-Matcher approach
- `extract_jd_keywords(jd_text)` β€” pulls keywords from the specific JD
- `jd_match_score(resume_text, jd_text)` β€” word-boundary regex matching (not substring)
- Final score: **70% JD match + 30% resume quality**
- Benchmark: EdTech JD β†’ 90%, SAP/ERP JD β†’ 53% (correctly differentiates)
**Files changed:** `src/ats_scorer.py` (full rewrite)
---
### Speed Optimization β€” 10-Model Parallel Pool
**Problem:** GLM 5.1 alone = 234s/job. 110 jobs = 6+ hours.
**Solution:** `ModelPool` with worker queue
- Phase 1 (keyword scoring): instant, no LLM
- Phase 2 (LLM assessment): 7 fast models compete for batches of 8 jobs
- Kimi K2.6 handles most work at ~5s/batch
- Wall clock for 110 jobs: ~3–5 minutes
**Files changed:** `src/model_pool.py`, `src/job_assessor.py`
---
### Added Models (cumulative)
| Model | API Key Env | Speed | Phase 2 |
|-------|------------|-------|---------|
| GLM-5.1 | NVIDIA_API_KEY | ~234s | No |
| Kimi-K2.6 | NVIDIA_API_KEY_3 | ~5s | Yes |
| Step-3.7-Flash | NVIDIA_API_KEY_8 | ~8-35s | Yes |
| Qwen3.5-397b | NVIDIA_API_KEY_7 | ~9s | Yes |
| Qwen3.5-122b-v2 | NVIDIA_API_KEY_7 | ~12s | Yes |
| GPT-OSS-120b | NVIDIA_API_KEY_5 | ~11s | Yes |
| Qwen3.5-122b | NVIDIA_API_KEY_4 | ~40s | Yes |
| DeepSeek-v4-Pro | NVIDIA_API_KEY_2 | ~42s | Yes |
| DeepSeek-v4-Flash | NVIDIA_API_KEY_6 | ~229s | No |
| MiniMax-M2.7 | NVIDIA_API_KEY_2 | ~908s | No |
---
### Odysseus Deep Research Engine
Integrated the [Odysseus IterResearch](https://github.com/pewdiepie-archdaemon/odysseus) engine for company research.
**Architecture:** Think β†’ Search β†’ Extract β†’ Synthesize loop
- DuckDuckGo search with Bing fallback
- 12h page content cache (`data/research_cache/`)
- GLM 5.1 for all LLM steps
- `asyncio.to_thread` + OpenAI SDK (not raw httpx) for proper timeout handling
**Files:** `src/research/deep_researcher.py`, `src/research/search.py`, `src/odysseus_llm_core.py`
---
### Google Sheets Integration
**Sheet columns:** Batch Date, Rank, Job Title, Company, Location, Platform, Salary, Experience, Relevance Score, ATS Before (%), ATS After (%), ATS Improvement, Resume Quality, Priority, Matching Skills, Missing Skills, AI Recommendation, Apply Link, Resume Link, Application Status, Date Applied, Notes
**Auth approach:** OAuth (user login via browser, token saved to `google_token.json`)
- Setup: `python connect_google.py`
- Required: Add `saitejatirunagari@gmail.com` as test user at https://console.cloud.google.com/apis/credentials/consent
**File:** `src/gsheets.py`
---
### PM-Only Filter
All scrapers enforce `BaseScraper.is_pm_role(title)` at scrape time:
- Title must contain "product"
- Must match PM patterns: product manager, product owner, APM, senior PM, etc.
- Blocked: engineer, developer, teacher, sales, marketing manager, project manager, data analyst, etc.
- Test result: 16/16 accuracy on mixed title set
**File:** `src/scrapers/base.py`
---
### Job Deduplication
SQLite store at `data/job_history.db`:
- `is_duplicate(url, days=30)` β€” skip jobs seen in last 30 days
- `bulk_mark_seen(jobs)` β€” handles both dict and `Job` dataclass objects
- Stats: `get_stats()`, housekeep: `clear_old_entries(days=90)`
**File:** `src/job_history.py`
---
### Bugs Fixed (Session 2)
| Bug | Fix |
|-----|-----|
| Kimi returns `' ["[7,6,8]"]'` (wrapped string) | `_parse_score_array()` unwraps `["[string]"]` format |
| `score_resume_against_jd` ImportError | Added backward-compat alias in `ats_scorer.py` |
| `bulk_mark_seen` AttributeError on Job dataclass | `isinstance(job, dict)` check + `getattr()` for dataclass |
| GLM timeout in research engine | Switched to OpenAI SDK via `asyncio.to_thread()`, timeout=300s |
| Windows `UnicodeEncodeError` on box-drawing chars | `sys.stdout = io.TextIOWrapper(encoding="utf-8", errors="replace")` |
| Google OAuth "Access blocked" (403) | Add email as test user in GCP OAuth consent screen |
---
## Session 1 β€” Initial Build
### Project Created
**Goal:** Automate PM job search β†’ AI assessment β†’ ATS resume β†’ Google Sheet.
**Stack chosen:**
- Scraping: requests + BeautifulSoup for LinkedIn; Playwright for Indeed/Glassdoor (JS-rendered)
- AI: NVIDIA API (OpenAI-compatible endpoint), starting with GLM 5.1
- Resume: pdfplumber (parse) + python-docx (generate DOCX)
- Storage: SQLite (dedup), gspread (Google Sheets), Google Drive API
- UI: Streamlit
---
### Scrapers Built
| Platform | Method | Status |
|----------|--------|--------|
| LinkedIn | requests + BeautifulSoup | βœ… Working |
| Indeed | Playwright (JS rendering) | βœ… Working |
| Glassdoor | Playwright | βœ… Working |
| Naukri | Attempted Playwright + requests | ❌ Blocked by Akamai (returns 406 / "Access Denied") |
**Key fixes during scraper development:**
- LinkedIn: company from `span[data-testid=company-name]`, title from `aria-label` (strip "full details of" prefix)
- Indeed: `div.job_seen_beacon` via BS4 on `page.content()` after `wait_until="networkidle"`
- Glassdoor: `li[data-jobid]` cards, `span[class*="compactEmployerName"]` for company
- Playwright sync_playwright conflict: two scrapers fighting over one context β†’ fixed by creating context per `search()` call
---
### Resume Parsing + Customization
- `ResumeParser` β€” pdfplumber extracts text from PDF
- `LLMClient` β€” GLM 5.1 extracts structured profile JSON + compact profile string
- `ResumeCustomizer` β€” iterative LLM optimizer:
1. LLM tailors resume to JD
2. Score it β†’ if < 95%, feed gap report back to LLM
3. Up to 3 attempts
4. Fallback: `_inject_missing_keywords()` to force 95%+
- Resume filename: `{Company}_{JobTitle}.docx` (no score in filename, per user request)
- Score stored in Google Sheet, not filename
---
### Streamlit UI
Four tabs:
1. **Search** β€” configure roles/locations, toggle platforms, run pipeline
2. **Results** β€” table view of all jobs with color-coded scores
3. **Job Details** β€” expand any job for full AI breakdown + resume download
4. **Deep Research** β€” Odysseus engine with quick-preset buttons from top jobs
Live progress via `_progress_q` queue + `st.rerun()` polling loop.
**File:** `ui.py`
---
## Pending (as of 2026-06-11)
| Task | Priority | Notes |
|------|----------|-------|
| Fix Google Drive upload `'Client' object has no attribute 'auth'` | High | gspread doesn't expose Drive auth directly |
| Fix LLM resume generation = 0 (GLM timeout) | High | Switch `ResumeCustomizer` to use Kimi/Step instead of GLM |
| Set `test_mode: False` in `config.py` | High | For full 100+ job production run |
| LLM-extracted JD keywords in ATS scoring | Medium | Use Kimi/Step to semantically extract required skills from each JD β†’ upgrade ATS from 7.5/10 to ~9/10 accuracy |
| Add `saitejatirunagari@gmail.com` as GCP test user | Done (user action) | https://console.cloud.google.com/apis/credentials/consent |