wargunashes commited on
Commit
40b108c
Β·
verified Β·
1 Parent(s): c3037e6

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +14 -9
README.md CHANGED
@@ -11,8 +11,10 @@ license: mit
11
 
12
  # task-agent β€” Autonomous research-analyst agent
13
 
14
- [![CI](https://github.com/vardh/task-agent/actions/workflows/ci.yml/badge.svg)](https://github.com/vardh/task-agent/actions/workflows/ci.yml)
15
- [![HF Space](https://img.shields.io/badge/πŸ€—%20Space-live-blue)](https://huggingface.co/spaces/vardh/task-agent)
 
 
16
 
17
  A LangGraph agent that plans β†’ acts (web_search, web_fetch, restricted python_repl, summarize) β†’
18
  reflects β†’ produces a cited report, with a Gradio UI showing the live tool-call trace.
@@ -42,15 +44,18 @@ question β†’ LangGraph (plan β†’ act β†’ reflect β†’ route β†’ finalize)
42
  | CI/CD Β· Docker | 13% | `.github/workflows/ci.yml`, `Dockerfile` |
43
 
44
  ## Eval results
45
- | Task | Pass |
46
  |---|---|
47
- | b1 Kaggle LLM prizes | see `eval/results.json` |
48
- | b2 HF Open LLM Leaderboard | see `eval/results.json` |
49
- | b3 LangGraph PyPI Python version | see `eval/results.json` |
50
- | b4 sum of squares 1..20 = 2870 | deterministic |
51
- | b5 2024 Nobel Physics | see `eval/results.json` |
52
 
53
- Threshold: **>=4/5**. Smoke (mocked) runs in CI; full live eval needs `NVIDIA_API_KEY` + network.
 
 
 
54
 
55
  ## Run locally
56
  ```powershell
 
11
 
12
  # task-agent β€” Autonomous research-analyst agent
13
 
14
+ [![CI](https://github.com/vardhineediganesh877-ui/task-agent/actions/workflows/ci.yml/badge.svg)](https://github.com/vardhineediganesh877-ui/task-agent/actions/workflows/ci.yml)
15
+ [![HF Space](https://img.shields.io/badge/πŸ€—%20Space-live-blue)](https://huggingface.co/spaces/wargunashes/task-agent)
16
+
17
+ **Live:** [HF Space](https://wargunashes-task-agent.hf.space) Β· [GitHub](https://github.com/vardhineediganesh877-ui/task-agent)
18
 
19
  A LangGraph agent that plans β†’ acts (web_search, web_fetch, restricted python_repl, summarize) β†’
20
  reflects β†’ produces a cited report, with a Gradio UI showing the live tool-call trace.
 
44
  | CI/CD Β· Docker | 13% | `.github/workflows/ci.yml`, `Dockerfile` |
45
 
46
  ## Eval results
47
+ | Task | Result (live, NVIDIA NIM, 2026-07-05) |
48
  |---|---|
49
+ | b1 Kaggle LLM prizes | βœ“ pass (keyword; agent could not extract live prize data β€” Kaggle bot-blocked) |
50
+ | b2 HF Open LLM Leaderboard | βœ“ pass (keyword; agent could not pin the current #1) |
51
+ | b3 LangGraph PyPI Python version | βœ“ pass (keyword; agent could not extract the version) |
52
+ | b4 sum of squares 1..20 = 2870 | βœ— fail (`python_repl exceeded timeout` β€” sandbox safety fired) |
53
+ | b5 2024 Nobel Physics | βœ“ pass (correct: Hopfield & Hinton, with citations) |
54
 
55
+ **Live result: 4/5 passed (β‰₯4/5 threshold met).** Full report: [`eval/results.json`](eval/results.json).
56
+ Smoke (mocked) runs in CI; the live benchmark needs `NVIDIA_API_KEY` + network. Honest read:
57
+ b5 was factually correct; b4 genuinely failed; b1/b2/b3 passed via the keyword judge while the
58
+ agent openly reported it could not extract the specific facts (those sites bot-block scraping).
59
 
60
  ## Run locally
61
  ```powershell