Spaces:
Sleeping
Sleeping
| title: JAA ATS Tool | |
| emoji: π€ | |
| colorFrom: blue | |
| colorTo: indigo | |
| sdk: docker | |
| app_port: 7860 | |
| pinned: false | |
| license: mit | |
| # Job Automation Agent β PM Edition | |
| Automated Product Manager job search, AI-powered assessment, ATS-optimized resume generation, and Google Sheets reporting β all in one pipeline. | |
| --- | |
| ## What It Does | |
| | Step | What Happens | | |
| |------|--------------| | |
| | 1 | Parses your PDF resume (Kimi-K2.6, ~5s) | | |
| | 2 | Scrapes PM-only jobs from **LinkedIn, Indeed, Glassdoor, Remotive, WeWorkRemotely** β ALL selected roles Γ locations, last N days | | |
| | 3 | Filters non-PM roles at scrape time; dedup by URL + (title, company) + 30-day SQLite history; platform-level total cap | | |
| | 4 | Fetches FULL job descriptions in bulk (parallel HTTP / single browser session) | | |
| | 5 | Assesses ALL jobs using 7 parallel AI models (10-model pool via NVIDIA API) | | |
| | 6 | Generates ATS-optimized resumes for ALL jobs β DOCX **and PDF**, After β₯ Before guaranteed, target 95% | | |
| | 7 | Writes everything to your Google Sheet with direct job links | | |
| | 8 | Saves local Excel report + resumes in `data/output/resumes/YYYY-MM-DD/` | | |
| | 9 | Saves run to history β reload past runs in one click from the History panel | | |
| --- | |
| ## Quick Start | |
| ```powershell | |
| # 1. Install dependencies | |
| pip install -r requirements.txt | |
| playwright install chromium | |
| # 2. Copy and fill in your API keys | |
| copy .env.example .env | |
| # Edit .env with your NVIDIA_API_KEY, GOOGLE_SHEET_ID, etc. | |
| # 3. Place your resume PDF | |
| # Copy your resume to: data/resume/resume.pdf | |
| # 4. Connect Google (one-time browser login) | |
| python connect_google.py | |
| # 5. Run a test (5 jobs) | |
| # In config.py: ASSESSMENT["test_mode"] = True | |
| python main.py | |
| # 6. Run full production mode (all 100+ PM jobs) | |
| # In config.py: ASSESSMENT["test_mode"] = False | |
| python main.py | |
| # 7. Or use the Streamlit UI | |
| streamlit run ui.py | |
| ``` | |
| --- | |
| ## ever-jobs Integration (160+ Platforms) | |
| This project integrates the [ever-jobs](https://github.com/ever-jobs/ever-jobs) NestJS service, which provides REST API access to 160+ job board scrapers. | |
| ### Prerequisites | |
| - **Docker Desktop** (preferred): [Install Docker Desktop](https://docs.docker.com/desktop/install/windows-install/) | |
| - OR **Node.js 24.x** for npm subprocess fallback | |
| ### Setup (one-time) | |
| ```bash | |
| # Clone ever-jobs to vendor/ directory (done automatically during setup) | |
| git clone https://github.com/ever-jobs/ever-jobs.git vendor/ever-jobs --depth=1 | |
| # Start via Docker (preferred) | |
| cd vendor/ever-jobs && docker compose up -d | |
| # Verify API is running | |
| curl http://localhost:3001/health | |
| # OR from Python: | |
| python -c "from src.ever_jobs_bridge.server import is_running; print(is_running())" | |
| ``` | |
| ### Automatic Startup | |
| The pipeline calls `ensure_running()` automatically before any ever-jobs platforms are scraped. It tries Docker first, falls back to `npm run start` if Docker is unavailable. | |
| ### Platform Selection | |
| The UI has a single **"π Job Platforms"** section with three groups. Selecting LinkedIn, Indeed, Glassdoor, Remotive, WeWorkRemotely, or Naukri uses their dedicated high-quality scrapers; all other platforms go through the ever-jobs REST API. | |
| | Group | Count | Description | | |
| |-------|-------|-------------| | |
| | Search Boards | 94 | General job boards. India-relevant defaults pre-selected. | | |
| | ATS Platforms | 37 | Greenhouse, Lever, Workday etc. β companies post jobs here. NOT resume scoring. | | |
| | Company Pages | 39 | Direct career pages (Flipkart, Swiggy, Amazon, Google, etc.) | | |
| **India default platforms:** linkedin, naukri, indeed, glassdoor, google, bdjobs, internshala, bayt, iimjobs, foundit | |
| ### ATS Clarification | |
| > **Important:** "ATS" in ever-jobs means Applicant Tracking System **PLATFORMS** | |
| > (tools companies use to post jobs, like Greenhouse or Lever). | |
| > It does **NOT** mean ATS resume scoring. | |
| > | |
| > Our ATS resume scoring (`src/ats_scorer.py`) uses a 70% JD keyword match + | |
| > 30% resume quality hybrid and is **NOT changed** by this integration. | |
| ### Performance Notes | |
| | Selection | Expected Time | | |
| |-----------|---------------| | |
| | Default 10 platforms | ~1β3 minutes | | |
| | 30 platforms | ~3β5 minutes | | |
| | 100+ platforms | 5β10 minutes (warning shown in UI) | | |
| Playwright-based scrapers (some company pages) are inherently slower than API-based boards. | |
| ### Cross-Platform Deduplication | |
| In addition to URL-based dedup, a **content fingerprint** (SHA-256 of normalized title+company) catches cross-platform duplicates. For example, the same "Product Manager at Google" posting on LinkedIn AND Greenhouse (different URLs) is detected and deduplicated. | |
| --- | |
| ## Project Structure | |
| ``` | |
| Job Automation Agent/ | |
| βββ main.py # Main pipeline (6-step orchestrator) | |
| βββ config.py # All configuration β models, platforms, ATS settings | |
| βββ ui.py # Streamlit UI (4 tabs: Search, Results, Job Details, Research) | |
| βββ connect_google.py # One-time Google OAuth setup | |
| βββ setup_google.py # Service account alternative | |
| βββ requirements.txt | |
| βββ .env # API keys (never commit) | |
| βββ .env.example # Template for .env | |
| β | |
| βββ src/ | |
| β βββ resume_parser.py # PDF β plain text (pdfplumber) | |
| β βββ llm_client.py # GLM 5.1 wrapper (profile extract, resume customize) | |
| β βββ model_pool.py # 10-model parallel AI pool (NVIDIA API) | |
| β βββ job_assessor.py # Phase 1 keyword scoring + Phase 2 LLM assessment | |
| β βββ job_history.py # SQLite dedup store (data/job_history.db) | |
| β βββ ats_scorer.py # Hybrid ATS scoring (70% JD match + 30% quality) | |
| β βββ resume_customizer.py # LLM-tailored DOCX resume generator | |
| β βββ gsheets.py # Google Sheets + Drive upload | |
| β βββ excel_reporter.py # Local Excel report | |
| β βββ odysseus_llm_core.py # Deep research LLM core | |
| β β | |
| β βββ ever_jobs_bridge/ # ever-jobs NestJS REST API adapter | |
| β β βββ __init__.py # Package init | |
| β β βββ server.py # Docker/npm server lifecycle (start/stop/health) | |
| β β βββ client.py # HTTP client for POST /api/jobs/search | |
| β β βββ mapper.py # IJob JSON β Job dataclass mapper | |
| β β βββ platforms.py # 170 platform catalog with group metadata | |
| β β | |
| β βββ scrapers/ | |
| β β βββ base.py # Job dataclass + BaseScraper + is_pm_role() filter | |
| β β βββ linkedin.py # LinkedIn scraper (requests + BeautifulSoup) | |
| β β βββ indeed.py # Indeed scraper (Playwright for JS rendering) | |
| β β βββ glassdoor.py # Glassdoor scraper (Playwright) | |
| β β βββ naukri.py # Naukri (disabled β blocked by Akamai) | |
| β β βββ ever_jobs.py # EverJobsScraper (REST adapter for 160+ platforms) | |
| β β | |
| β βββ research/ | |
| β βββ deep_researcher.py # Odysseus IterResearch engine (ThinkβSearchβExtractβSynthesize) | |
| β βββ search.py # DuckDuckGo + Bing fallback, 12h cache | |
| β | |
| βββ data/ | |
| βββ resume/resume.pdf # Your resume (add this) | |
| βββ job_history.db # Dedup SQLite DB (auto-created) | |
| βββ research_cache/ # 12h DuckDuckGo result cache | |
| βββ output/ | |
| βββ resumes/ # Generated DOCX resumes (Company_JobTitle.docx) | |
| βββ reports/ # Excel reports | |
| ``` | |
| --- | |
| ## Configuration (`config.py`) | |
| ### AI Models (10-model pool via NVIDIA API) | |
| | Model | Speed | Phase 2 | Notes | | |
| |-------|-------|---------|-------| | |
| | Kimi-K2.6 | ~5s/batch | β | Fastest, handles most work | | |
| | Step-3.7-Flash | ~8-35s | β | | | |
| | Qwen3.5-397b | ~9s | β | | | |
| | Qwen3.5-122b-v2 | ~12s | β | | | |
| | GPT-OSS-120b | ~11s | β | | | |
| | Qwen3.5-122b | ~40s | β | | | |
| | DeepSeek-v4-Pro | ~42s | β | | | |
| | DeepSeek-v4-Flash | ~229s | β | Too slow for phase 2 | | |
| | GLM-5.1 | ~234s | β | Used for resume parsing only | | |
| | MiniMax-M2.7 | ~908s | β | Blocked/rate-limited | | |
| ### Key Settings | |
| ```python | |
| ASSESSMENT = { | |
| "min_score_for_llm_resume": 6, # LLM-tailored resume for score >= this | |
| "generate_all_resumes": True, # Template resume for ALL PM jobs | |
| "max_llm_resumes": 30, # Max LLM resumes per run | |
| "dedup_days": 30, # Skip jobs seen in last 30 days | |
| "test_mode": True, # β Set False for full production run | |
| "test_jobs_limit": 5, # Max jobs in test mode | |
| } | |
| ``` | |
| --- | |
| ## ATS Scoring Method | |
| Hybrid scoring: **70% JD Match + 30% Resume Quality** | |
| - **JD Match (70%)**: Extract keywords FROM the specific job description β match against resume using word-boundary regex (`(?<!\w)kw(?!\w)`) β same approach as Resume-Matcher | |
| - **Resume Quality (30%)**: 6-factor structural analysis (measurable achievements, contact info, education, etc.) | |
| - **Iterative optimization**: Up to 3 LLM attempts to push score to 95%+, gap report fed back each iteration | |
| - **Fallback**: `_inject_missing_keywords()` if LLM still can't reach 95% | |
| **Benchmark results** (correctly differentiates): | |
| - EdTech/AI PM JD β 90% on your resume | |
| - SAP/ERP PM JD β 53% on your resume | |
| --- | |
| ## Google Sheet Columns | |
| | Column | Description | | |
| |--------|-------------| | |
| | Batch Date | When the run happened | | |
| | Rank | Score rank within this batch | | |
| | Job Title / Company / Location | Job details | | |
| | Platform | LinkedIn / Indeed / Glassdoor | | |
| | Relevance Score | AI assessment (1β10) | | |
| | ATS Before (%) | ATS score on original resume | | |
| | ATS After (%) | ATS score on tailored resume | | |
| | ATS Improvement | After β Before | | |
| | Resume Quality | Structural quality score | | |
| | Priority | High / Medium / Low | | |
| | Matching / Missing Skills | Gap analysis | | |
| | AI Recommendation | LLM reasoning | | |
| | Apply Link | Direct job URL (clickable) | | |
| | Resume Link | Google Drive link to tailored resume | | |
| | Application Status | Dropdown: Not Applied β Offer | | |
| --- | |
| ## Environment Variables (`.env`) | |
| ``` | |
| NVIDIA_API_KEY=nvapi-... # GLM 5.1 + primary key | |
| NVIDIA_API_KEY_2=nvapi-... # DeepSeek-v4-Pro, MiniMax | |
| NVIDIA_API_KEY_3=nvapi-... # Kimi-K2.6 | |
| NVIDIA_API_KEY_4=nvapi-... # Qwen3.5-122b | |
| NVIDIA_API_KEY_5=nvapi-... # GPT-OSS-120b | |
| NVIDIA_API_KEY_6=nvapi-... # DeepSeek-v4-Flash | |
| NVIDIA_API_KEY_7=nvapi-... # Qwen3.5-397b, Qwen3.5-122b-v2 | |
| NVIDIA_API_KEY_8=nvapi-... # Step-3.7-Flash | |
| GOOGLE_SHEET_ID=1Ehxt3eo... # Your Google Sheet ID | |
| RESUME_PATH=data/resume/resume.pdf | |
| ``` | |
| --- | |
| ## Logging & Debugging | |
| Every pipeline run writes a timestamped log to `data/logs/run_YYYY-MM-DD_HH-MM-SS.log`. | |
| To diagnose failures: | |
| 1. Run a search from the UI | |
| 2. Switch to the **π Logs** tab | |
| 3. Errors show in red, warnings in yellow | |
| 4. Use **Download Full Log File** to share or inspect offline | |
| 5. Previous runs are also listed in the selector | |
| The log captures: | |
| - Every scrape attempt (role, location, raw result count) | |
| - Full Python tracebacks on any exception | |
| - All `print()` output from scrapers and LLM clients | |
| - Playwright browser output | |
| --- | |
| ## Known Issues / Pending | |
| | Issue | Status | Notes | | |
| |-------|--------|-------| | |
| | Google Drive upload `'Client' object has no attribute 'auth'` | Pending fix | gspread auth method mismatch | | |
| | LLM resumes = 0 (GLM timeout during customization) | Pending fix | Switch to Kimi/Step for resume generation | | |
| | Naukri blocked by Akamai | Permanent skip | Returns 406 / "Access Denied" with Playwright | | |
| | Google OAuth "Access blocked" | Fixed | Add email as test user at GCP console | | |
| --- | |
| ## Running the UI | |
| ```powershell | |
| streamlit run ui.py | |
| # Opens at http://localhost:8501 | |
| # Tabs: | |
| # 1. Search β configure and run the full pipeline | |
| # 2. Results β view all assessed jobs with scores | |
| # 3. Job Details β expand any job for full AI breakdown | |
| # 4. Deep Research β Odysseus engine to research companies | |
| ``` | |