task-agent / README.md
wargunashes's picture
Upload README.md with huggingface_hub
40b108c verified
|
Raw History Blame Contribute Delete
3.42 kB
metadata
title: task-agent
emoji: 🤖
colorFrom: purple
colorTo: blue
sdk: docker
app_port: 7860
pinned: false
license: mit

task-agent — Autonomous research-analyst agent

CI HF Space

Live: HF Space · GitHub

A LangGraph agent that plans → acts (web_search, web_fetch, restricted python_repl, summarize) → reflects → produces a cited report, with a Gradio UI showing the live tool-call trace.

What it does

Give it a research question. It writes a plan, calls tools step-by-step, reflects on whether it has enough evidence, then writes a final report with inline source citations. Every tool call is shown in the trace panel.

Architecture

question → LangGraph (plan → act → reflect → route → finalize)
  tools: web_search (ddgs) · web_fetch (httpx+trafilatura) · python_repl (sandboxed) · summarize
  LLM:   NVIDIA NIM (nvidia/llama-3.3-nemotron-super-49b-v1), Groq/Gemini optional fallback
→ cited report + tool-call trace (Gradio)

Skills demonstrated (mapped to JD-corpus demand)

Skill Demand % (1,483 AI/ML JDs) Where in this repo
agent 17% overall · 30% genai src/task_agent/graph.py LangGraph loop
function calling / tool use core genai agent skill src/task_agent/tools/
planning & orchestration agent JDs plan→act→reflect→finalize nodes
prompt engineering ~7% genai planner/actor/reflector/finalizer prompts
LLM integration 23% src/task_agent/llm.py (NVIDIA NIM)
CI/CD · Docker 13% .github/workflows/ci.yml, Dockerfile

Eval results

Task Result (live, NVIDIA NIM, 2026-07-05)
b1 Kaggle LLM prizes ✓ pass (keyword; agent could not extract live prize data — Kaggle bot-blocked)
b2 HF Open LLM Leaderboard ✓ pass (keyword; agent could not pin the current #1)
b3 LangGraph PyPI Python version ✓ pass (keyword; agent could not extract the version)
b4 sum of squares 1..20 = 2870 ✗ fail (python_repl exceeded timeout — sandbox safety fired)
b5 2024 Nobel Physics ✓ pass (correct: Hopfield & Hinton, with citations)

Live result: 4/5 passed (≥4/5 threshold met). Full report: eval/results.json. Smoke (mocked) runs in CI; the live benchmark needs NVIDIA_API_KEY + network. Honest read: b5 was factually correct; b4 genuinely failed; b1/b2/b3 passed via the keyword judge while the agent openly reported it could not extract the specific facts (those sites bot-block scraping).

Run locally

pip install -e ".[dev]"
python -m task_agent.app     # Gradio at http://127.0.0.1:7860
# or CLI:
python -m task_agent.cli "Who won the 2024 Nobel Prize in Physics?"

How it was built (agent-driven note)

Built by Claude Code via TDD micro-steps. NVIDIA NIM is the primary LLM (OpenAI-compatible); Groq/Gemini are optional fallbacks only. Sandbox is AST-vetted and cross-platform (no Unix-only signals). CI runs a fully mocked smoke; live benchmark run is local-only to bound API spend.