--- title: task-agent emoji: 🤖 colorFrom: purple colorTo: blue sdk: docker app_port: 7860 pinned: false license: mit --- # task-agent — Autonomous research-analyst agent [![CI](https://github.com/vardhineediganesh877-ui/task-agent/actions/workflows/ci.yml/badge.svg)](https://github.com/vardhineediganesh877-ui/task-agent/actions/workflows/ci.yml) [![HF Space](https://img.shields.io/badge/🤗%20Space-live-blue)](https://huggingface.co/spaces/wargunashes/task-agent) **Live:** [HF Space](https://wargunashes-task-agent.hf.space) · [GitHub](https://github.com/vardhineediganesh877-ui/task-agent) A LangGraph agent that plans → acts (web_search, web_fetch, restricted python_repl, summarize) → reflects → produces a cited report, with a Gradio UI showing the live tool-call trace. ## What it does Give it a research question. It writes a plan, calls tools step-by-step, reflects on whether it has enough evidence, then writes a final report with inline source citations. Every tool call is shown in the trace panel. ## Architecture ``` question → LangGraph (plan → act → reflect → route → finalize) tools: web_search (ddgs) · web_fetch (httpx+trafilatura) · python_repl (sandboxed) · summarize LLM: NVIDIA NIM (nvidia/llama-3.3-nemotron-super-49b-v1), Groq/Gemini optional fallback → cited report + tool-call trace (Gradio) ``` ## Skills demonstrated (mapped to JD-corpus demand) | Skill | Demand % (1,483 AI/ML JDs) | Where in this repo | |---|---|---| | agent | 17% overall · 30% genai | `src/task_agent/graph.py` LangGraph loop | | function calling / tool use | core genai agent skill | `src/task_agent/tools/` | | planning & orchestration | agent JDs | plan→act→reflect→finalize nodes | | prompt engineering | ~7% genai | planner/actor/reflector/finalizer prompts | | LLM integration | 23% | `src/task_agent/llm.py` (NVIDIA NIM) | | CI/CD · Docker | 13% | `.github/workflows/ci.yml`, `Dockerfile` | ## Eval results | Task | Result (live, NVIDIA NIM, 2026-07-05) | |---|---| | b1 Kaggle LLM prizes | ✓ pass (keyword; agent could not extract live prize data — Kaggle bot-blocked) | | b2 HF Open LLM Leaderboard | ✓ pass (keyword; agent could not pin the current #1) | | b3 LangGraph PyPI Python version | ✓ pass (keyword; agent could not extract the version) | | b4 sum of squares 1..20 = 2870 | ✗ fail (`python_repl exceeded timeout` — sandbox safety fired) | | b5 2024 Nobel Physics | ✓ pass (correct: Hopfield & Hinton, with citations) | **Live result: 4/5 passed (≥4/5 threshold met).** Full report: [`eval/results.json`](eval/results.json). Smoke (mocked) runs in CI; the live benchmark needs `NVIDIA_API_KEY` + network. Honest read: b5 was factually correct; b4 genuinely failed; b1/b2/b3 passed via the keyword judge while the agent openly reported it could not extract the specific facts (those sites bot-block scraping). ## Run locally ```powershell pip install -e ".[dev]" python -m task_agent.app # Gradio at http://127.0.0.1:7860 # or CLI: python -m task_agent.cli "Who won the 2024 Nobel Prize in Physics?" ``` ## How it was built (agent-driven note) Built by Claude Code via TDD micro-steps. NVIDIA NIM is the primary LLM (OpenAI-compatible); Groq/Gemini are optional fallbacks only. Sandbox is AST-vetted and cross-platform (no Unix-only signals). CI runs a fully mocked smoke; live benchmark run is local-only to bound API spend.