rakeshmadasaniai commited on
Commit ·
1e2fc7c
1
Parent(s): e1ffae6
Publish autonomy evaluation audit and refreshed evaluation report
Browse files
01-rag-system/README.md
CHANGED
|
@@ -136,6 +136,7 @@ Supporting scripts:
|
|
| 136 |
- `summarize_eval_sets.py`
|
| 137 |
|
| 138 |
Latest committed result snapshots live in [`evaluation/results`](./evaluation/results).
|
|
|
|
| 139 |
|
| 140 |
### Latest committed summaries
|
| 141 |
|
|
|
|
| 136 |
- `summarize_eval_sets.py`
|
| 137 |
|
| 138 |
Latest committed result snapshots live in [`evaluation/results`](./evaluation/results).
|
| 139 |
+
Autonomy audit for current release is tracked in [`../AUTONOMY_EVALUATION.md`](../AUTONOMY_EVALUATION.md).
|
| 140 |
|
| 141 |
### Latest committed summaries
|
| 142 |
|
01-rag-system/evaluation/reports/latest_portfolio_report.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
| 1 |
# Portfolio Evaluation Report
|
| 2 |
|
| 3 |
-
Generated: 2026-05-
|
| 4 |
|
| 5 |
## Domain Pack Snapshot
|
| 6 |
|
|
|
|
| 1 |
# Portfolio Evaluation Report
|
| 2 |
|
| 3 |
+
Generated: 2026-05-11 21:04:44 UTC
|
| 4 |
|
| 5 |
## Domain Pack Snapshot
|
| 6 |
|
AUTONOMY_EVALUATION.md
ADDED
|
@@ -0,0 +1,56 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Autonomy Evaluation (Current Release)
|
| 2 |
+
|
| 3 |
+
Last updated: 2026-05-11 (America/Chicago)
|
| 4 |
+
|
| 5 |
+
This document evaluates whether the app behaves as an autonomous agent, based on the current code and committed test/evaluation artifacts.
|
| 6 |
+
|
| 7 |
+
## Verdict
|
| 8 |
+
|
| 9 |
+
- **Autonomous agent mode exists and is functional:** Yes
|
| 10 |
+
- **Fully autonomous production system (strict definition):** Partial
|
| 11 |
+
|
| 12 |
+
## What is implemented now
|
| 13 |
+
|
| 14 |
+
1. **Tool-calling autonomous runtime**
|
| 15 |
+
- `01-rag-system/core/agentic_runtime.py` uses an OpenAI tools loop with iterative tool execution and observations.
|
| 16 |
+
|
| 17 |
+
2. **Decision preflight**
|
| 18 |
+
- Explicit preflight routing for scam/structuring/sanctions/math signals before generic clarification.
|
| 19 |
+
|
| 20 |
+
3. **In-loop verification + self-correction**
|
| 21 |
+
- Verification tool runs before finalization.
|
| 22 |
+
- If verification fails (`needs_retry`), rewrite happens inside the same loop.
|
| 23 |
+
|
| 24 |
+
4. **Autonomous supervisor mode**
|
| 25 |
+
- `Autonomous Max` mode continues execution with explicit assumptions when clarifications would otherwise block progress.
|
| 26 |
+
|
| 27 |
+
5. **Persistent autonomous operations**
|
| 28 |
+
- Background task queue and policy audit log are persisted:
|
| 29 |
+
- `01-rag-system/data/autonomous_queue.json`
|
| 30 |
+
- `01-rag-system/data/autonomous_audit_log.jsonl`
|
| 31 |
+
|
| 32 |
+
6. **Regression tests**
|
| 33 |
+
- `01-rag-system/tests/test_agentic_decision_engine.py` validates core decision behavior.
|
| 34 |
+
|
| 35 |
+
## Current evidence
|
| 36 |
+
|
| 37 |
+
- Local decision-engine regression tests: **6/6 passing**
|
| 38 |
+
- Portfolio report regenerated from committed evaluation summaries:
|
| 39 |
+
- `01-rag-system/evaluation/reports/latest_portfolio_report.md`
|
| 40 |
+
|
| 41 |
+
## Remaining gap to “fully autonomous” (strict enterprise bar)
|
| 42 |
+
|
| 43 |
+
To claim strict full autonomy in enterprise environments, these are still recommended:
|
| 44 |
+
|
| 45 |
+
1. Scheduled autonomous execution service outside Streamlit session lifecycle.
|
| 46 |
+
2. Explicit approval workflows for high-risk policy decisions (human gate with signed trace).
|
| 47 |
+
3. Stronger SLA telemetry (p50/p95/p99 latency, tool failure rates, retry budgets).
|
| 48 |
+
4. Expanded test matrix for all languages/input types in CI (voice, PDF, DOCX, image).
|
| 49 |
+
|
| 50 |
+
## Honest label to use publicly
|
| 51 |
+
|
| 52 |
+
Use:
|
| 53 |
+
- **“Autonomous tool-calling banking AI agent with policy-aware execution and audit logging.”**
|
| 54 |
+
|
| 55 |
+
Avoid overclaiming:
|
| 56 |
+
- **“Fully autonomous AGI”**.
|
README.md
CHANGED
|
@@ -30,6 +30,7 @@ I did not want this to be a one-screen chatbot demo. I wanted it to behave like
|
|
| 30 |
I maintain a concrete week-by-week delivery plan in:
|
| 31 |
|
| 32 |
- [`WORLDCLASS_WEEK1_TO_WEEK6.md`](WORLDCLASS_WEEK1_TO_WEEK6.md)
|
|
|
|
| 33 |
|
| 34 |
This is the operating plan used to move the project from strong prototype quality to production-grade architecture, reliability, governance, and distribution.
|
| 35 |
|
|
|
|
| 30 |
I maintain a concrete week-by-week delivery plan in:
|
| 31 |
|
| 32 |
- [`WORLDCLASS_WEEK1_TO_WEEK6.md`](WORLDCLASS_WEEK1_TO_WEEK6.md)
|
| 33 |
+
- [`AUTONOMY_EVALUATION.md`](AUTONOMY_EVALUATION.md)
|
| 34 |
|
| 35 |
This is the operating plan used to move the project from strong prototype quality to production-grade architecture, reliability, governance, and distribution.
|
| 36 |
|