# task-agent — system design ## Problem Recruiters want evidence of agentic systems: planning, tool use, reflection, and grounded output. This project is an autonomous research-analyst agent that takes a question, plans, calls real tools, reflects on evidence, and returns a cited report with a visible tool-call trace. ## Constraints - $0 infra (HF Spaces CPU, free LLM tier). - NVIDIA NIM (OpenAI-compatible) primary LLM; Groq/Gemini optional fallback only. - No local GPU. Windows + PowerShell dev host (sandbox must not rely on Unix `signal.SIGALRM`). ## Architecture ``` User question ↓ LangGraph StateGraph: plan → act → reflect → (route: act | finalize) → END ↓ Tools (act node chooses one per step via JSON): • web_search — DuckDuckGo (ddgs) • web_fetch — httpx + trafilatura main-text extraction • python_repl — AST-vetted restricted sandbox, thread-timeout, no imports/I/O • summarize — NIM-backed abstractive summary ↓ NVIDIA NIM (nvidia/llama-3.3-nemotron-super-49b-v1) drives planner/actor/reflector/finalizer ↓ Final cited report + live tool-call trace (Gradio) ``` ## Trade-offs - **Rule-based judge vs LLM-as-judge:** benchmark uses deterministic keyword/source checks so pass/fail is reproducible and free (no judge LLM spend). - **Single-tool-per-step actor** (JSON) over parallel tool arrays: simpler control flow, easier to trace; capped at `MAX_STEPS=6` to bound cost. - **Cross-platform sandbox:** AST safety check + `ThreadPoolExecutor` timeout (works on Windows; `signal.SIGALRM` is Unix-only). Threads can't be hard-killed, so the AST gate is the real safety barrier — execution time is a liveness guard, not a security boundary. - **Smoke mocks** the agent entirely in CI (zero network/LLM spend); live run behind local exec. ## Eval strategy 5-task benchmark (`eval/benchmark_tasks.json`) over stable, verifiable facts (sum of squares 1..20 = 2870; 2024 Nobel Physics = Hopfield & Hinton; PyPI metadata; Kaggle/HF leaderboard URLs). Pass threshold ≥ 4/5. Judge = required keywords + optional source-URL presence.