Spaces:
Sleeping
Sleeping
| # Project History β Job Automation Agent | |
| A running log of everything built, fixed, and changed. Most recent first. | |
| --- | |
| ## 2026-06-13 β Unified Platform Selector + ATS + HTML Rendering Fixes | |
| ### Changes | |
| - **Unified platform selector**: Merged the 6 legacy checkboxes ("π Job Platforms") and the grouped ever-jobs selector ("π ever-jobs Platforms") into a single "π Job Platforms" section. One place to search all 170 platforms. Selecting LinkedIn/Indeed/Glassdoor/Remotive/WeWorkRemotely/Naukri still routes to their dedicated high-quality scrapers; everything else goes through EverJobsScraper. | |
| - **ATS min_score default**: Changed slider default from 6 to 1 β LLM resumes now generated for ALL jobs regardless of score. | |
| - **HTML rendering fix**: Switched all 5 `st.markdown(..., unsafe_allow_html=True)` calls to `st.html()` β fixes raw `<span>`/`<a>` tags showing as plain text in job cards (Streamlit 1.45+ regression). | |
| ### Modified Files | |
| - `ui.py` β removed 6 legacy checkboxes, renamed section label, updated platforms_cfg, updated pipeline routing to use unified `all_platforms` key | |
| --- | |
| ## 2026-06-13 β Phase 1: ever-jobs Integration (160+ Platforms) | |
| ### New Features | |
| - **160+ job platforms** via ever-jobs REST API integration (was 5 platforms) | |
| - **Grouped platform selector** in UI: Search Boards / ATS Platforms / Company Pages with st.multiselect search | |
| - **India-focused defaults**: 10 platforms pre-selected (LinkedIn, Naukri, Indeed, Glassdoor, Google, BDJobs, Internshala, Bayt, IIMJobs, Foundit) | |
| - **Content fingerprint dedup**: SHA-256 of (title+company) catches cross-platform duplicates where same job appears on LinkedIn AND Greenhouse with different URLs | |
| - **Performance warning**: UI shows warning when >30 platforms selected | |
| ### New Files | |
| - `src/ever_jobs_bridge/__init__.py` β package init | |
| - `src/ever_jobs_bridge/server.py` β Docker/npm server lifecycle (start/stop/health) | |
| - `src/ever_jobs_bridge/client.py` β HTTP client for POST /api/jobs/search | |
| - `src/ever_jobs_bridge/mapper.py` β IJob JSON β Job dataclass field mapper | |
| - `src/ever_jobs_bridge/platforms.py` β 170 platform catalog with group metadata | |
| - `src/scrapers/ever_jobs.py` β EverJobsScraper extending BaseScraper | |
| - `vendor/ever-jobs/` β ever-jobs NestJS monorepo (cloned, gitignored) | |
| ### Modified Files | |
| - `src/job_history.py` β added content_fp column + is_duplicate_by_content() function | |
| - `config.py` β added EVER_JOBS config block | |
| - `ui.py` β grouped platform selector + EverJobsScraper pipeline wiring + ever_jobs step | |
| - `requirements.txt` β added rapidfuzz>=3.0 | |
| - `.gitignore` β added vendor/ | |
| ### R3 ATS Finding (Definitive) | |
| ever-jobs "ATS" = Applicant Tracking System platforms that companies use to POST jobs | |
| (Greenhouse, Lever, Workday). This is NOT resume scoring. | |
| Our `src/ats_scorer.py` (70% JD keyword match + 30% resume quality) is the correct | |
| resume ATS scoring system and is UNCHANGED. No modifications to ats_scorer.py are needed. | |
| ### Backward Compatibility | |
| All existing scrapers (LinkedIn, Indeed, Glassdoor, Remotive, WeWorkRemotely) are UNTOUCHED. | |
| Pipeline flow is unchanged β ever-jobs is an additive parallel path. | |
| --- | |
| ## Session 10 β 2026-06-13 | |
| ### New: 2 additional job platforms (Remotive + We Work Remotely) | |
| - **`src/scrapers/remotive.py`** β Remotive.io public JSON API. No auth needed. | |
| Fetches WFH/remote PM jobs globally (India-eligible: "Worldwide" / APAC filter). | |
| - **`src/scrapers/weworkremotely.py`** β We Work Remotely RSS feed scraper. | |
| Free-to-scrape, good volume of remote PM roles. | |
| - Both expose `get_details_bulk()` (no-op, descriptions come with the listing). | |
| - Both appear as checkboxes in the new UI; step-skip if unchecked. | |
| ### Fixed: max_resumes slider removed β all jobs now get a resume | |
| Previously `max_resumes` slider (default 15) silently capped LLM resumes even | |
| when 30β40 jobs were fetched. Fixed by passing `max_llm_resumes=len(assessed_jobs)` | |
| (effectively no cap). Every eligible job now gets an LLM-tailored resume. | |
| ### Fixed: platform cap is now total-per-platform, not per-query | |
| Old code applied `max_results=N` per roleΓlocation query. With 3 roles Γ 3 locations | |
| you could get 9 Γ 15 = 135 from one platform β far more than the user intended. | |
| New code: the outer loop breaks once `platform_jobs` reaches `max_jobs_per_platform`, | |
| and the per-query `max_results` is set to `remaining = cap - len(platform_jobs)`. | |
| ### Fixed: Google Sheets error messages are now informative | |
| - `FileNotFoundError` (no credentials) now emits a clear "run setup_google.py" hint | |
| - Full error text (up to 120 chars) logged to the live UI log, not just the file log | |
| - A "Google Sheet status" indicator (β/β ) shown in the Configure section before run | |
| ### New: run history (save + load past runs) | |
| - **`src/run_history.py`** β saves each completed run as JSON in | |
| `data/output/run_history/run_YYYY-MM-DD_HH-MM-SS.json`. | |
| Summary fields stored without jobs for fast listing; full jobs on load. | |
| - History is auto-saved at the end of every pipeline run. | |
| - UI "Load" button restores any past run's results to the active session without | |
| rerunning the pipeline. | |
| ### New: complete UI redesign (ui.py) | |
| - **No sidebar** β all controls now live inline in the main area. | |
| - **History panel** β top-right "π History" button opens a panel listing all | |
| past runs with stats (jobs, high-priority count, ATS before/after). Click "Load" | |
| to restore any run. | |
| - **Configure section** β expandable card with resume upload, roles, locations, | |
| platform checkboxes, days, max-per-platform, and min score. Google Sheet | |
| status shown inline. | |
| - **Start button** β centered, prominent, full-width. | |
| - **Step timeline** β CSS grid layout (auto-fill columns), fits all platforms. | |
| - **Results tab β job cards** β top 10 shown as visual cards (title, company, | |
| ATS before/after, salary, apply link). Switch to "Full Table" for all jobs. | |
| - **Download fix** β zip now contains only the current run's date subfolder (not | |
| all historical date folders). Eliminates the "90 files for 30 jobs" confusion | |
| (per run: 30 DOCX + 30 PDF = 60 files as expected). | |
| - **Metrics row** β Total | High | Medium | LLM Resumes | PDFs | Avg ATS After. | |
| - Welcome state shown when no results are loaded yet. | |
| ### Fixed: test_mode β False in config.py | |
| Was accidentally left `True`, capping the pipeline at 10 jobs per test run. | |
| --- | |
| ## Session 9 β 2026-06-13 | |
| ### Fixed: UI stuck at "0% β Startingβ¦" while pipeline ran fine in background | |
| **Symptom:** Click Start β UI shows 0% and all steps "Waitingβ¦" forever, but the | |
| console/logs show the pipeline scraping, assessing 41 jobs, and generating | |
| resumes at 91β94% ATS. Users clicked Start again thinking it was dead β duplicate | |
| pipeline threads (Thread-8 + Thread-17 in the logs). | |
| **Root cause:** `_progress_q = queue.Queue()` was created at MODULE level in | |
| ui.py with a comment claiming module globals survive reruns. They do NOT β | |
| Streamlit re-executes the entry script top-to-bottom on EVERY rerun, creating a | |
| brand-new empty Queue each time. The background thread kept writing progress to | |
| the original queue; the UI drain loop polled the new empty one. Nothing ever | |
| arrived. | |
| **Fix (ui.py):** | |
| - Queue now lives in `st.session_state["progress_q"]` β the only store that | |
| survives reruns within a session | |
| - `run_pipeline` receives the queue as an explicit default arg (`_q=_progress_q`) | |
| and shadows the module helpers, so the thread always writes to the queue the | |
| drain loop reads β even across reruns and multiple sessions | |
| - `st.session_state["current_log_file"]` was being set FROM the background | |
| thread (the "missing ScriptRunContext" warning, silently broken) β now sent | |
| through the queue as a `("logfile", path)` message handled by the drain loop | |
| **Verified with Streamlit AppTest:** queue identity preserved across reruns; | |
| clicked Start in the test harness β UI received 7 log messages, step cards | |
| updated (resume β β profile β β linkedin β³), progress bar at 15%. | |
| **Files changed:** `ui.py`, `HISTORY.md` | |
| --- | |
| ## Session 8 β 2026-06-12 | |
| ### Major performance + quality overhaul: parallel resumes, PDF output, full JD fetching | |
| **Root causes of "taking lot of time, not going forward":** | |
| 1. LLM resumes generated ONE at a time (50β150s each Γ 30 = up to an hour, UI frozen) | |
| 2. Indeed launched a full Chromium browser PER job description (~10s overhead each) | |
| 3. Glassdoor NEVER fetched descriptions (no detail method existed) | |
| 4. LinkedIn `job_id` regex broken β LinkedIn switched to slug URLs | |
| (`/jobs/view/title-at-company-4423634421`), so ALL detail fetches 404'd β no JDs | |
| 5. UI capped search to 3 roles Γ 2 locations | |
| **Fixes:** | |
| - `src/resume_customizer.py` β LLM resumes now generated IN PARALLEL via | |
| ThreadPoolExecutor (6 workers, round-robin across phase2 model API keys). | |
| Per-resume `progress_cb` streams live status to the UI. | |
| - `src/scrapers/linkedin.py` β fixed job_id extraction (slug URLs); new | |
| `get_details_bulk()` fetches ALL descriptions with 4 parallel HTTP workers | |
| - `src/scrapers/indeed.py` β new `get_details_bulk()`: ONE browser session for | |
| all job descriptions instead of one browser per job | |
| - `src/scrapers/glassdoor.py` β new `get_details_bulk()` with Cloudflare-challenge | |
| wait + JSON-LD JobPosting parsing (Glassdoor still intermittent β bot-hostile) | |
| - `ui.py` β searches ALL selected roles Γ locations (caps removed); cross-platform | |
| dedup by (title, company) in addition to URL; live per-resume progress | |
| **ATS quality fixes (tailored resumes were sometimes scoring LOWER than original):** | |
| - `src/llm_client.py` β validates LLM customization (summary >50 chars, β₯5 skills), | |
| retries once, unwraps JSON arrays, max_tokens 3000β4000 | |
| - `resume_customizer.py` β optimization loop now: scores with same extra_kw as | |
| final report Β· skips empty customizations Β· retries fall back to Kimi Β· rewrites | |
| BEST attempt to disk (was keeping last) Β· GUARANTEE: if LLM result scores below | |
| the original resume, ships keyword-injected template instead (After β₯ Before always) | |
| - `_inject_missing_keywords()` rewritten β now injects the ACTUAL missing JD | |
| keywords (was injecting generic PM keywords that didn't move the JD-match score) | |
| **PDF output (new):** | |
| - `src/pdf_writer.py` β DOCXβPDF: one Word COM session per batch on Windows | |
| (perfect fidelity), reportlab re-render fallback on Linux/HF Spaces | |
| - Every resume now saved as both `.docx` and `.pdf` in `data/output/resumes/YYYY-MM-DD/` | |
| - UI: PDF + DOCX download buttons per job; zip download includes PDFs | |
| - `requirements.txt`: + reportlab, docx2pdf (win32 only) | |
| **Files changed:** `src/pdf_writer.py` (new), `src/resume_customizer.py`, | |
| `src/llm_client.py`, `src/scrapers/linkedin.py`, `src/scrapers/indeed.py`, | |
| `src/scrapers/glassdoor.py`, `ui.py`, `requirements.txt`, `README.md`, `HISTORY.md` | |
| --- | |
| ## Session 7 β 2026-06-12 | |
| ### File-based logging system + Logs tab in UI | |
| **Problem:** Pipeline was failing on HF Spaces with no way to see why. Queue-based live log only showed last 30 messages and swallowed full tracebacks. | |
| **What was built:** | |
| **`src/app_logger.py`** β New centralized logger: | |
| - Writes every run to `data/logs/run_YYYY-MM-DD_HH-MM-SS.log` | |
| - Captures ALL Python logging output (INFO, WARNING, ERROR, DEBUG) | |
| - Redirects stdout/stderr via `_TeeStream` so `print()` and Playwright output are also captured | |
| - In-memory ring buffer (500 lines) for UI access without file I/O | |
| - `list_log_files()` returns all previous runs, newest first | |
| **`ui.py`** changes: | |
| - New **π Logs** tab (5th tab) | |
| - Color-coded viewer: errors=red, warnings=yellow, INFO done=green, info=blue | |
| - Slider to show 50β500 lines | |
| - Toggle to show/hide DEBUG lines | |
| - Auto-refresh every 2s while pipeline is running | |
| - Download button for raw `.log` file | |
| - Previous run selector to load any past log | |
| - Error/warning counts in footer | |
| - Pipeline thread now calls `app_logger.setup()` at start β creates timestamped log file | |
| - Every scrape attempt logged with role + location + raw result count | |
| - Full tracebacks on scrape errors (`logging.error(..., traceback)`) | |
| - Fatal pipeline exceptions logged in full, not truncated to 400 chars | |
| - `current_log_file` added to session state defaults | |
| **`Dockerfile`** β Added `data/logs` to `mkdir -p` list | |
| **Files changed:** `src/app_logger.py` (new), `ui.py`, `Dockerfile`, `HISTORY.md`, `README.md` | |
| --- | |
| ## Session 6 β 2026-06-11 | |
| ### GitHub push + Hugging Face Spaces deployment prep | |
| **Code pushed to GitHub:** https://github.com/saitejatiru/JAA-ATS-Tool | |
| **HF Spaces files added:** | |
| - `README.md` β prepended YAML frontmatter (`sdk: streamlit`, `app_file: ui.py`) | |
| - `packages.txt` β Chromium system dependencies for Playwright on Linux | |
| - `.gitignore` β excludes secrets (`google_token.json`, `.env`, resumes, output data) | |
| - `.env.example` β documents all 9 NVIDIA API keys + Google Sheet ID | |
| - `requirements.txt` β added `gspread`, `google-auth`, `google-auth-oauthlib`, `google-api-python-client` | |
| **`ui.py` changes for HF Spaces:** | |
| - Playwright install: `@st.cache_resource` function installs Chromium once per server lifetime | |
| - Google credentials bootstrap: reads `GOOGLE_CREDENTIALS_JSON` env var and writes to `google_credentials.json` on startup | |
| **Files changed:** `README.md`, `requirements.txt`, `packages.txt`, `.gitignore`, `.env.example`, `ui.py` | |
| --- | |
| ## Session 5 β 2026-06-11 | |
| ### ATS Before/After in Excel + Verbose Resume Error Logging | |
| **Excel reporter fixed:** | |
| - Added `ATS Before (%)`, `ATS After (%)`, `ATS Improvement` columns to all sheets (was completely missing) | |
| - Column order: Relevance Score β ATS Before β ATS After β ATS Improvement β Skills Match β β¦ | |
| - `_pct()` helper: shows `"45%"` or `"β"` for null; improvement shows `"+37pp"` or `"β"` | |
| - Column indices for score badge (9), URL hyperlink (23), priority color (15) updated to match new order | |
| **Resume error visibility:** | |
| - Added explicit `tqdm.write()` on success: `"β LLM resume: Google β ATS 45% β 82% (+37pp)"` | |
| - Added `traceback.format_exc()` on failure so exact error is visible in the terminal | |
| - Fallback ATS scoring (original resume score) always runs on failure so sheet never shows blank | |
| **Confirmed working (run completed 2026-06-11 11:16):** | |
| - 7 LLM-tailored + 2 template resumes generated in `data/output/resumes/2026-06-11/` | |
| - Google Sheet updated with all 10 jobs | |
| - Files: Google_Product Manager I Ads.docx, Instagram, Workday, Giga, Denave, Tessera, Latinem | |
| **Files changed:** `src/excel_reporter.py`, `src/resume_customizer.py` | |
| --- | |
| ## Session 4 β 2026-06-11 | |
| ### ATS Before/After Fix + Best Resume Prompt | |
| **ATS Before/After not showing β root causes fixed:** | |
| 1. `score_resume()` was calling Kimi AGAIN (via `fast_model_cfg`) during ATS scoring β after already using Kimi for 9 resume generations, rate limits caused silent failures and blank scores. Fixed: removed `fast_model_cfg` from scoring calls; use pre-extracted keywords from assessment phase only. | |
| 2. On resume generation failure, `ats_score_before/after` was never set at all. Fixed: fallback block now always computes and stores ATS scores even if DOCX generation fails. | |
| **Best ATS resume β prompt redesigned:** | |
| - Old prompt: generic instructions, 1500 char JD limit, 2000 token output | |
| - New prompt: | |
| - Explicit mandatory keyword list with instruction "MUST include ALL of these" | |
| - Rules enforce: exact JD language mirroring, action verbs on every bullet, quantified metrics required | |
| - JD limit raised to 2000 chars, resume to 2500 chars | |
| - Output tokens raised to 3000 (room for full detailed resume) | |
| - 15 core competencies (was 12) | |
| - More specific bullet format: "β’ Led X resulting in Y% improvement" | |
| **Profile extraction speed fix:** | |
| - Step 2 was blocked on GLM 5.1 (~234s). Now tries Kimi-K2.6 (~5s) first via `extract_profile_summary_fast(cfg, ...)` with fallback to GLM. | |
| - Added `LLMClient.extract_profile_summary_fast(cfg, resume_text)` method. | |
| **Files changed:** `src/llm_client.py`, `src/resume_customizer.py`, `main.py` | |
| --- | |
| ## Session 3 β 2026-06-11 | |
| ### Streamlit UI Fixes + LLM Resume Root-Cause Fix | |
| **4 issues addressed:** | |
| | Issue | Fix | | |
| |-------|-----| | |
| | LLM resumes = 0 | Root cause: `ATSScorer` class imported but never existed β silent `ImportError`. Fixed by replacing with `score_resume()` function. Also fixed `PM_DOMAIN_KEYWORDS` β `PM_BASE_KEYWORDS + PM_TOOLS` | | |
| | Fast model for resume generation | Added `LLMClient._call_with_cfg()` + `customize_resume_fast(cfg, ...)`. Now uses Kimi-K2.6 (~5s) instead of GLM (~234s) | | |
| | Date-based local resume folders | Resumes now save to `data/output/resumes/YYYY-MM-DD/`. No more Google Drive upload | | |
| | Sheet headers missing | `gsheets.py` now detects missing header row and inserts at row 1 using `ws.insert_row()` even when data already exists | | |
| | Test limit | 5 β 10 jobs | | |
| **Streamlit UI updated:** | |
| - Fixed `customize_for_jobs()` parameter mismatch (`min_score` β `min_score_for_llm`, `max_count` β `max_llm_resumes`) | |
| - Resume zip download now scans all date subfolders (`Path.rglob("*.docx")`) | |
| - Results table now shows **ATS Before, ATS After, ATS Gain** columns | |
| - Job Details tab shows ATS before/after inline | |
| - `fast_model_cfg` wired into UI pipeline (Kimi-K2.6 for LLM keywords + resume tailoring) | |
| **To launch UI:** | |
| ```powershell | |
| streamlit run ui.py | |
| # Opens at http://localhost:8501 | |
| ``` | |
| --- | |
| ## Session 2 β 2026-06-11 | |
| ### Test Run Completed Successfully β | |
| **Results:** | |
| - LinkedIn 60 + Indeed 18 + Glassdoor 13 jobs scraped (capped to 5 in test mode) | |
| - Assessment: **16 seconds** for 5 jobs (Kimi K2.6, single batch) | |
| - Top job: Associate Product Manager (Adtech) at MakeMyTrip β Score 8/10 | |
| - Google Sheet updated: https://docs.google.com/spreadsheets/d/1Ehxt3eortehbtySdtgcSrMhCqmxIMUAmvRqSkII0HJk/edit | |
| - Excel saved: `data/output/reports/job_report.xlsx` | |
| - 5 jobs marked in dedup store (SQLite) β won't reappear next run | |
| **Bugs found during test run:** | |
| 1. `bulk_mark_seen` AttributeError β `Job` dataclass doesn't have `.get()`. Fixed with `isinstance(job, dict)` + `getattr()`. | |
| 2. Drive upload: `'Client' object has no attribute 'auth'` β gspread doesn't expose Drive API directly. **Still pending fix.** | |
| 3. LLM resumes = 0 β resume customization calling GLM (234s), timing out silently. **Still pending fix** (need to switch to Kimi/Step). | |
| --- | |
| ### ATS Scoring β Rebuilt from Scratch | |
| **Problem:** Original ATS scored resume quality (structural), not job-description match. A generic resume scored the same for any job. | |
| **Solution:** Resume-Matcher approach | |
| - `extract_jd_keywords(jd_text)` β pulls keywords from the specific JD | |
| - `jd_match_score(resume_text, jd_text)` β word-boundary regex matching (not substring) | |
| - Final score: **70% JD match + 30% resume quality** | |
| - Benchmark: EdTech JD β 90%, SAP/ERP JD β 53% (correctly differentiates) | |
| **Files changed:** `src/ats_scorer.py` (full rewrite) | |
| --- | |
| ### Speed Optimization β 10-Model Parallel Pool | |
| **Problem:** GLM 5.1 alone = 234s/job. 110 jobs = 6+ hours. | |
| **Solution:** `ModelPool` with worker queue | |
| - Phase 1 (keyword scoring): instant, no LLM | |
| - Phase 2 (LLM assessment): 7 fast models compete for batches of 8 jobs | |
| - Kimi K2.6 handles most work at ~5s/batch | |
| - Wall clock for 110 jobs: ~3β5 minutes | |
| **Files changed:** `src/model_pool.py`, `src/job_assessor.py` | |
| --- | |
| ### Added Models (cumulative) | |
| | Model | API Key Env | Speed | Phase 2 | | |
| |-------|------------|-------|---------| | |
| | GLM-5.1 | NVIDIA_API_KEY | ~234s | No | | |
| | Kimi-K2.6 | NVIDIA_API_KEY_3 | ~5s | Yes | | |
| | Step-3.7-Flash | NVIDIA_API_KEY_8 | ~8-35s | Yes | | |
| | Qwen3.5-397b | NVIDIA_API_KEY_7 | ~9s | Yes | | |
| | Qwen3.5-122b-v2 | NVIDIA_API_KEY_7 | ~12s | Yes | | |
| | GPT-OSS-120b | NVIDIA_API_KEY_5 | ~11s | Yes | | |
| | Qwen3.5-122b | NVIDIA_API_KEY_4 | ~40s | Yes | | |
| | DeepSeek-v4-Pro | NVIDIA_API_KEY_2 | ~42s | Yes | | |
| | DeepSeek-v4-Flash | NVIDIA_API_KEY_6 | ~229s | No | | |
| | MiniMax-M2.7 | NVIDIA_API_KEY_2 | ~908s | No | | |
| --- | |
| ### Odysseus Deep Research Engine | |
| Integrated the [Odysseus IterResearch](https://github.com/pewdiepie-archdaemon/odysseus) engine for company research. | |
| **Architecture:** Think β Search β Extract β Synthesize loop | |
| - DuckDuckGo search with Bing fallback | |
| - 12h page content cache (`data/research_cache/`) | |
| - GLM 5.1 for all LLM steps | |
| - `asyncio.to_thread` + OpenAI SDK (not raw httpx) for proper timeout handling | |
| **Files:** `src/research/deep_researcher.py`, `src/research/search.py`, `src/odysseus_llm_core.py` | |
| --- | |
| ### Google Sheets Integration | |
| **Sheet columns:** Batch Date, Rank, Job Title, Company, Location, Platform, Salary, Experience, Relevance Score, ATS Before (%), ATS After (%), ATS Improvement, Resume Quality, Priority, Matching Skills, Missing Skills, AI Recommendation, Apply Link, Resume Link, Application Status, Date Applied, Notes | |
| **Auth approach:** OAuth (user login via browser, token saved to `google_token.json`) | |
| - Setup: `python connect_google.py` | |
| - Required: Add `saitejatirunagari@gmail.com` as test user at https://console.cloud.google.com/apis/credentials/consent | |
| **File:** `src/gsheets.py` | |
| --- | |
| ### PM-Only Filter | |
| All scrapers enforce `BaseScraper.is_pm_role(title)` at scrape time: | |
| - Title must contain "product" | |
| - Must match PM patterns: product manager, product owner, APM, senior PM, etc. | |
| - Blocked: engineer, developer, teacher, sales, marketing manager, project manager, data analyst, etc. | |
| - Test result: 16/16 accuracy on mixed title set | |
| **File:** `src/scrapers/base.py` | |
| --- | |
| ### Job Deduplication | |
| SQLite store at `data/job_history.db`: | |
| - `is_duplicate(url, days=30)` β skip jobs seen in last 30 days | |
| - `bulk_mark_seen(jobs)` β handles both dict and `Job` dataclass objects | |
| - Stats: `get_stats()`, housekeep: `clear_old_entries(days=90)` | |
| **File:** `src/job_history.py` | |
| --- | |
| ### Bugs Fixed (Session 2) | |
| | Bug | Fix | | |
| |-----|-----| | |
| | Kimi returns `' ["[7,6,8]"]'` (wrapped string) | `_parse_score_array()` unwraps `["[string]"]` format | | |
| | `score_resume_against_jd` ImportError | Added backward-compat alias in `ats_scorer.py` | | |
| | `bulk_mark_seen` AttributeError on Job dataclass | `isinstance(job, dict)` check + `getattr()` for dataclass | | |
| | GLM timeout in research engine | Switched to OpenAI SDK via `asyncio.to_thread()`, timeout=300s | | |
| | Windows `UnicodeEncodeError` on box-drawing chars | `sys.stdout = io.TextIOWrapper(encoding="utf-8", errors="replace")` | | |
| | Google OAuth "Access blocked" (403) | Add email as test user in GCP OAuth consent screen | | |
| --- | |
| ## Session 1 β Initial Build | |
| ### Project Created | |
| **Goal:** Automate PM job search β AI assessment β ATS resume β Google Sheet. | |
| **Stack chosen:** | |
| - Scraping: requests + BeautifulSoup for LinkedIn; Playwright for Indeed/Glassdoor (JS-rendered) | |
| - AI: NVIDIA API (OpenAI-compatible endpoint), starting with GLM 5.1 | |
| - Resume: pdfplumber (parse) + python-docx (generate DOCX) | |
| - Storage: SQLite (dedup), gspread (Google Sheets), Google Drive API | |
| - UI: Streamlit | |
| --- | |
| ### Scrapers Built | |
| | Platform | Method | Status | | |
| |----------|--------|--------| | |
| | LinkedIn | requests + BeautifulSoup | β Working | | |
| | Indeed | Playwright (JS rendering) | β Working | | |
| | Glassdoor | Playwright | β Working | | |
| | Naukri | Attempted Playwright + requests | β Blocked by Akamai (returns 406 / "Access Denied") | | |
| **Key fixes during scraper development:** | |
| - LinkedIn: company from `span[data-testid=company-name]`, title from `aria-label` (strip "full details of" prefix) | |
| - Indeed: `div.job_seen_beacon` via BS4 on `page.content()` after `wait_until="networkidle"` | |
| - Glassdoor: `li[data-jobid]` cards, `span[class*="compactEmployerName"]` for company | |
| - Playwright sync_playwright conflict: two scrapers fighting over one context β fixed by creating context per `search()` call | |
| --- | |
| ### Resume Parsing + Customization | |
| - `ResumeParser` β pdfplumber extracts text from PDF | |
| - `LLMClient` β GLM 5.1 extracts structured profile JSON + compact profile string | |
| - `ResumeCustomizer` β iterative LLM optimizer: | |
| 1. LLM tailors resume to JD | |
| 2. Score it β if < 95%, feed gap report back to LLM | |
| 3. Up to 3 attempts | |
| 4. Fallback: `_inject_missing_keywords()` to force 95%+ | |
| - Resume filename: `{Company}_{JobTitle}.docx` (no score in filename, per user request) | |
| - Score stored in Google Sheet, not filename | |
| --- | |
| ### Streamlit UI | |
| Four tabs: | |
| 1. **Search** β configure roles/locations, toggle platforms, run pipeline | |
| 2. **Results** β table view of all jobs with color-coded scores | |
| 3. **Job Details** β expand any job for full AI breakdown + resume download | |
| 4. **Deep Research** β Odysseus engine with quick-preset buttons from top jobs | |
| Live progress via `_progress_q` queue + `st.rerun()` polling loop. | |
| **File:** `ui.py` | |
| --- | |
| ## Pending (as of 2026-06-11) | |
| | Task | Priority | Notes | | |
| |------|----------|-------| | |
| | Fix Google Drive upload `'Client' object has no attribute 'auth'` | High | gspread doesn't expose Drive auth directly | | |
| | Fix LLM resume generation = 0 (GLM timeout) | High | Switch `ResumeCustomizer` to use Kimi/Step instead of GLM | | |
| | Set `test_mode: False` in `config.py` | High | For full 100+ job production run | | |
| | LLM-extracted JD keywords in ATS scoring | Medium | Use Kimi/Step to semantically extract required skills from each JD β upgrade ATS from 7.5/10 to ~9/10 accuracy | | |
| | Add `saitejatirunagari@gmail.com` as GCP test user | Done (user action) | https://console.cloud.google.com/apis/credentials/consent | | |