Spaces:
Sleeping
Sleeping
rakeshmadasaniai commited on
Commit ·
7d48dea
1
Parent(s): ca99372
Polish AI agent portfolio documentation
Browse files- .env.example +13 -0
- .gitignore +67 -2
- 01-rag-system/README.md +73 -170
- 02-qa-dataset/README.md +71 -53
- 03-qlora-finetuning/README.md +58 -54
- 04-conversational-memory/README.md +48 -195
- CONTRIBUTING.md +62 -0
- EVALUATION.md +77 -0
- README.md +135 -240
- ROADMAP.md +49 -0
- SYSTEM_DESIGN.md +118 -0
- requirements.txt +5 -1
.env.example
ADDED
|
@@ -0,0 +1,13 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
OPENAI_API_KEY=
|
| 2 |
+
HF_TOKEN=
|
| 3 |
+
MODEL_MODE=OpenAI
|
| 4 |
+
TOP_K=4
|
| 5 |
+
TEMPERATURE=0.2
|
| 6 |
+
|
| 7 |
+
# Optional runtime configuration
|
| 8 |
+
OPENAI_MODEL=gpt-4o-mini
|
| 9 |
+
OPENAI_STT_MODEL=gpt-4o-mini-transcribe
|
| 10 |
+
OPENAI_TTS_MODEL=gpt-4o-mini-tts
|
| 11 |
+
OPENAI_TTS_VOICE=alloy
|
| 12 |
+
FINETUNED_MODEL_ID=RakeshMadasani/banking-finance-mistral-qlora
|
| 13 |
+
FINETUNED_ENDPOINT_URL=
|
.gitignore
CHANGED
|
@@ -1,7 +1,72 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
.venv/
|
|
|
|
|
|
|
| 2 |
**/.venv/
|
| 3 |
-
**/
|
| 4 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 5 |
01-rag-system/screenshots/*.png
|
| 6 |
01-rag-system/screenshots/*.jpg
|
| 7 |
01-rag-system/screenshots/*.jpeg
|
|
|
|
| 1 |
+
# Python
|
| 2 |
+
__pycache__/
|
| 3 |
+
**/__pycache__/
|
| 4 |
+
*.py[cod]
|
| 5 |
+
*.pyo
|
| 6 |
+
*.pyd
|
| 7 |
+
.Python
|
| 8 |
+
*.egg-info/
|
| 9 |
+
.pytest_cache/
|
| 10 |
+
.mypy_cache/
|
| 11 |
+
.ruff_cache/
|
| 12 |
+
.coverage
|
| 13 |
+
htmlcov/
|
| 14 |
+
|
| 15 |
+
# Virtual environments
|
| 16 |
.venv/
|
| 17 |
+
venv/
|
| 18 |
+
env/
|
| 19 |
**/.venv/
|
| 20 |
+
**/venv/
|
| 21 |
+
|
| 22 |
+
# Environment and secrets
|
| 23 |
+
.env
|
| 24 |
+
.env.*
|
| 25 |
+
!.env.example
|
| 26 |
+
**/.env
|
| 27 |
+
**/.env.*
|
| 28 |
+
!**/.env.example
|
| 29 |
+
.streamlit/secrets.toml
|
| 30 |
+
**/.streamlit/secrets.toml
|
| 31 |
+
|
| 32 |
+
# Streamlit and local app state
|
| 33 |
+
.streamlit/
|
| 34 |
+
**/.streamlit/
|
| 35 |
+
*.log
|
| 36 |
+
logs/
|
| 37 |
+
|
| 38 |
+
# Local model artifacts and generated indexes
|
| 39 |
+
models/
|
| 40 |
+
checkpoints/
|
| 41 |
+
outputs/
|
| 42 |
+
wandb/
|
| 43 |
+
runs/
|
| 44 |
+
mlruns/
|
| 45 |
+
*.safetensors
|
| 46 |
+
*.bin
|
| 47 |
+
*.pt
|
| 48 |
+
*.pth
|
| 49 |
+
*.onnx
|
| 50 |
+
*.gguf
|
| 51 |
+
**/faiss_index/
|
| 52 |
+
**/vectorstore/
|
| 53 |
+
**/*.faiss
|
| 54 |
+
**/*.pkl
|
| 55 |
+
|
| 56 |
+
# Data artifacts that should not be committed accidentally
|
| 57 |
+
*.parquet
|
| 58 |
+
*.arrow
|
| 59 |
+
*.sqlite
|
| 60 |
+
*.db
|
| 61 |
+
*.duckdb
|
| 62 |
+
|
| 63 |
+
# OS and editor files
|
| 64 |
+
.DS_Store
|
| 65 |
+
Thumbs.db
|
| 66 |
+
.vscode/
|
| 67 |
+
.idea/
|
| 68 |
+
|
| 69 |
+
# Large local screenshots; curated tracked screenshots can be force-added if needed
|
| 70 |
01-rag-system/screenshots/*.png
|
| 71 |
01-rag-system/screenshots/*.jpg
|
| 72 |
01-rag-system/screenshots/*.jpeg
|
01-rag-system/README.md
CHANGED
|
@@ -1,194 +1,97 @@
|
|
| 1 |
-
#
|
| 2 |
|
| 3 |
-
Banking & Finance
|
| 4 |
|
| 5 |
-
##
|
| 6 |
|
| 7 |
-
|
| 8 |
|
| 9 |
-
##
|
| 10 |
-
|
| 11 |
-
This is the product layer of the broader portfolio. The goal was not just to make a banking chatbot answer questions. The goal was to make it behave like a product someone could open, test, trust, and discuss seriously:
|
| 12 |
-
|
| 13 |
-
- grounded answers instead of free-floating generation
|
| 14 |
-
- visible source support
|
| 15 |
-
- multiple model modes with clear routing behavior
|
| 16 |
-
- uploads for real user documents
|
| 17 |
-
- multilingual interaction
|
| 18 |
-
- reproducible evaluation, not just screenshots
|
| 19 |
-
|
| 20 |
-
## Where The Code Is
|
| 21 |
-
|
| 22 |
-
The live Hugging Face Space is only the deployment target. The actual product implementation is committed here in this project:
|
| 23 |
-
|
| 24 |
-
- [`core`](./core)
|
| 25 |
-
retrieval orchestration, runtime flow, prompts, and shared utilities
|
| 26 |
-
- [`features`](./features)
|
| 27 |
-
UI rendering, uploads, read-aloud, answer formatting, and user interaction
|
| 28 |
-
- [`models`](./models)
|
| 29 |
-
OpenAI mode, Fine-Tuned mode, and Auto routing logic
|
| 30 |
-
|
| 31 |
-
That is important because I wanted the repo to stand on its own as a real product codebase, not just point outward to a demo URL.
|
| 32 |
-
|
| 33 |
-
## What The User Can Do
|
| 34 |
-
|
| 35 |
-
- ask banking, AML, KYC, FDIC, Basel III, RBI, and compliance questions
|
| 36 |
-
- switch between `OpenAI`, `Fine-Tuned`, and `Auto` modes
|
| 37 |
-
- upload PDF, DOCX, TXT, and image files
|
| 38 |
-
- inspect retrieved sources under each answer
|
| 39 |
-
- use read-aloud on the final response
|
| 40 |
-
- test multilingual questions
|
| 41 |
-
|
| 42 |
-
## Why The Product Is Structured This Way
|
| 43 |
-
|
| 44 |
-
Trust was the main design constraint. For finance and compliance questions, a polished answer alone is not enough. The product needs to show where the answer came from and make its behavior explainable.
|
| 45 |
-
|
| 46 |
-
That is why the app is built around:
|
| 47 |
-
|
| 48 |
-
- shared retrieval before generation
|
| 49 |
-
- source cards and chunk previews
|
| 50 |
-
- model-mode transparency
|
| 51 |
-
- latency and confidence visibility
|
| 52 |
-
- evaluation packs committed in the repository
|
| 53 |
-
|
| 54 |
-
## Model Modes
|
| 55 |
-
|
| 56 |
-
### OpenAI
|
| 57 |
-
|
| 58 |
-
This is the strongest general-purpose answer path and the most stable baseline for live testing.
|
| 59 |
-
|
| 60 |
-
### Fine-Tuned
|
| 61 |
-
|
| 62 |
-
This uses the banking-domain Mistral adapter. It is valuable when the hosted path is configured and when lower-latency or domain-style responses are desirable.
|
| 63 |
-
|
| 64 |
-
### Auto
|
| 65 |
-
|
| 66 |
-
Auto retrieves once, evaluates candidate answer paths, and selects the winner. That makes the routing logic easier to reason about than a hidden black-box switch.
|
| 67 |
-
|
| 68 |
-
## Product Architecture
|
| 69 |
|
| 70 |
```mermaid
|
| 71 |
-
flowchart
|
| 72 |
-
A["
|
| 73 |
-
B --> C["
|
| 74 |
-
C --> D["
|
| 75 |
-
D --> E["
|
| 76 |
-
E --> F["
|
| 77 |
-
|
| 78 |
-
|
| 79 |
-
F --> I["
|
| 80 |
-
|
| 81 |
-
|
| 82 |
-
|
|
|
|
|
|
|
|
|
|
| 83 |
```
|
| 84 |
|
| 85 |
-
##
|
| 86 |
-
|
| 87 |
-
1. The app loads curated banking knowledge files and any uploaded user documents.
|
| 88 |
-
2. Documents are chunked and embedded.
|
| 89 |
-
3. FAISS retrieves the most relevant context for the question.
|
| 90 |
-
4. The selected model mode answers from that shared context.
|
| 91 |
-
5. The UI renders the answer together with:
|
| 92 |
-
- mode
|
| 93 |
-
- latency
|
| 94 |
-
- chunk count
|
| 95 |
-
- source cards
|
| 96 |
-
- confidence label
|
| 97 |
-
|
| 98 |
-
## Product Walkthrough
|
| 99 |
-
|
| 100 |
-
### A clean first impression for the product
|
| 101 |
-
|
| 102 |
-
This is the opening experience of the Banking & Finance Copilot: the stable sidebar, the mode selector, the welcome guidance, and multilingual starter questions that make the product feel usable from the first click.
|
| 103 |
-
|
| 104 |
-

|
| 105 |
-
|
| 106 |
-
### A grounded English answer that feels concise and useful
|
| 107 |
-
|
| 108 |
-
This example shows the assistant answering a KYC question in English with a direct explanation, short supporting bullets, visible latency, and a retrieved source card underneath the answer.
|
| 109 |
-
|
| 110 |
-

|
| 111 |
-
|
| 112 |
-
### The same product experience working in Telugu
|
| 113 |
|
| 114 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 115 |
|
| 116 |
-
|
| 117 |
|
| 118 |
-
|
| 119 |
-
|
| 120 |
-
|
| 121 |
-
|
| 122 |
-
|
| 123 |
-
|
| 124 |
-
## Evaluation
|
| 125 |
-
|
| 126 |
-
The [`evaluation`](./evaluation) folder includes two larger committed evaluation packs:
|
| 127 |
-
|
| 128 |
-
- `evaluation_queries.md`
|
| 129 |
-
120 domain-specific prompts across OpenAI, Fine-Tuned, and Auto
|
| 130 |
-
- `evaluation_multilingual.md`
|
| 131 |
-
120 multilingual prompts across the same three modes
|
| 132 |
-
|
| 133 |
-
Supporting scripts:
|
| 134 |
-
|
| 135 |
-
- `run_eval_sets.py`
|
| 136 |
-
- `summarize_eval_sets.py`
|
| 137 |
-
|
| 138 |
-
Latest committed result snapshots live in [`evaluation/results`](./evaluation/results).
|
| 139 |
-
Autonomy audit for current release is tracked in [`../AUTONOMY_EVALUATION.md`](../AUTONOMY_EVALUATION.md).
|
| 140 |
-
|
| 141 |
-
### Latest committed summaries
|
| 142 |
-
|
| 143 |
-
| Evaluation set | Total prompts | Available evaluated rows | Average latency | Median latency |
|
| 144 |
-
|---|---:|---:|---:|---:|
|
| 145 |
-
| Domain set | 120 | 80 | 2037.0 ms | 2036.0 ms |
|
| 146 |
-
| Multilingual set | 120 | 80 | 2031.8 ms | 2031.5 ms |
|
| 147 |
-
|
| 148 |
-
### Reading the snapshot correctly
|
| 149 |
-
|
| 150 |
-
Those numbers are the committed run snapshot, not a made-up "best case" table:
|
| 151 |
|
| 152 |
-
|
| 153 |
-
- 80 rows were available in the committed export
|
| 154 |
-
- the missing rows reflect backend availability in that local run, not missing evaluation logic
|
| 155 |
|
| 156 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 157 |
|
| 158 |
-
##
|
| 159 |
|
| 160 |
-
|
| 161 |
|
| 162 |
-
-
|
| 163 |
-
-
|
|
|
|
|
|
|
| 164 |
|
| 165 |
-
|
| 166 |
|
| 167 |
-
|
| 168 |
-
|
| 169 |
-
|
|
|
|
|
|
|
| 170 |
|
| 171 |
-
##
|
| 172 |
|
| 173 |
-
|
| 174 |
-
OPENAI_API_KEY=your_api_key_here
|
| 175 |
-
OPENAI_MODEL=gpt-4o-mini
|
| 176 |
-
OPENAI_STT_MODEL=gpt-4o-mini-transcribe
|
| 177 |
-
OPENAI_TTS_MODEL=gpt-4o-mini-tts
|
| 178 |
-
OPENAI_TTS_VOICE=alloy
|
| 179 |
-
FINETUNED_MODEL_ID=RakeshMadasani/banking-finance-mistral-qlora
|
| 180 |
-
FINETUNED_ENDPOINT_URL=
|
| 181 |
-
HF_TOKEN=your_hugging_face_token
|
| 182 |
-
```
|
| 183 |
|
| 184 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 185 |
|
| 186 |
-
|
| 187 |
-
streamlit run app.py
|
| 188 |
-
```
|
| 189 |
|
| 190 |
-
##
|
| 191 |
|
| 192 |
-
-
|
| 193 |
-
-
|
| 194 |
-
-
|
|
|
|
|
|
| 1 |
+
# AI Agent Runtime
|
| 2 |
|
| 3 |
+
This folder contains the live deployed Banking & Finance AI Agent runtime. It is still named `01-rag-system` for Hugging Face deployment compatibility, but its role in the system is the agent runtime: Streamlit UI, retrieval, model routing, agentic workflows, uploads, voice controls, source cards, and evaluation.
|
| 4 |
|
| 5 |
+
## Purpose
|
| 6 |
|
| 7 |
+
Provide a product-quality AI workflow interface for banking, compliance, and financial knowledge tasks. The runtime prioritizes grounded answers, clear confidence signals, source visibility, multilingual support, and stable user interaction.
|
| 8 |
|
| 9 |
+
## Architecture
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 10 |
|
| 11 |
```mermaid
|
| 12 |
+
flowchart TD
|
| 13 |
+
A["User input"] --> B["Streamlit product runtime"]
|
| 14 |
+
B --> C["Upload / voice / text handling"]
|
| 15 |
+
C --> D["Chunking and embeddings"]
|
| 16 |
+
D --> E["FAISS dense retrieval"]
|
| 17 |
+
E --> F["Grounded context"]
|
| 18 |
+
F --> G["OpenAI mode"]
|
| 19 |
+
F --> H["Fine-Tuned mode"]
|
| 20 |
+
F --> I["Auto mode"]
|
| 21 |
+
F --> J["Agentic / Autonomous modes"]
|
| 22 |
+
G --> K["Answer renderer"]
|
| 23 |
+
H --> K
|
| 24 |
+
I --> K
|
| 25 |
+
J --> K
|
| 26 |
+
K --> L["Sources, confidence, latency, actions, read aloud"]
|
| 27 |
```
|
| 28 |
|
| 29 |
+
## Key Files
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 30 |
|
| 31 |
+
| File or folder | Role |
|
| 32 |
+
|---|---|
|
| 33 |
+
| `app.py` | Streamlit entry point used by the live Hugging Face Space. |
|
| 34 |
+
| `core/product_runtime.py` | Main product orchestration, mode selection, retrieval calls, session state, and response handling. |
|
| 35 |
+
| `core/agentic_runtime.py` | Agentic/autonomous workflow implementation and tool-style reasoning layer. |
|
| 36 |
+
| `core/retriever.py` | Shared context retrieval over runtime indexes. |
|
| 37 |
+
| `core/vector_store.py` | FAISS vector store construction. |
|
| 38 |
+
| `features/product_ui.py` | Premium UI cards, sidebar, metrics, and answer rendering. |
|
| 39 |
+
| `features/voice_input.py` / `features/voice_output.py` | Speech-to-text and text-to-speech integration paths. |
|
| 40 |
+
| `models/openai_mode.py` | OpenAI answer path. |
|
| 41 |
+
| `models/finetuned_mode.py` | Fine-tuned model endpoint path. |
|
| 42 |
+
| `models/auto_router.py` | Candidate scoring and automatic model selection. |
|
| 43 |
+
| `evaluation/` | Domain and multilingual evaluation packs, runners, summaries, and reports. |
|
| 44 |
|
| 45 |
+
## How To Run
|
| 46 |
|
| 47 |
+
```bash
|
| 48 |
+
cd 01-rag-system
|
| 49 |
+
pip install -r requirements.txt
|
| 50 |
+
streamlit run app.py
|
| 51 |
+
```
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 52 |
|
| 53 |
+
Required environment:
|
|
|
|
|
|
|
| 54 |
|
| 55 |
+
```env
|
| 56 |
+
OPENAI_API_KEY=
|
| 57 |
+
HF_TOKEN=
|
| 58 |
+
OPENAI_MODEL=gpt-4o-mini
|
| 59 |
+
FINETUNED_MODEL_ID=RakeshMadasani/banking-finance-mistral-qlora
|
| 60 |
+
```
|
| 61 |
|
| 62 |
+
## Inputs And Outputs
|
| 63 |
|
| 64 |
+
Inputs:
|
| 65 |
|
| 66 |
+
- User questions in English or supported multilingual prompts.
|
| 67 |
+
- PDF, DOCX, TXT, and image-oriented upload workflows.
|
| 68 |
+
- Voice input where browser/runtime support is available.
|
| 69 |
+
- Mode selection across OpenAI, Fine-Tuned, Auto, Agentic Workspace, and Autonomous Max.
|
| 70 |
|
| 71 |
+
Outputs:
|
| 72 |
|
| 73 |
+
- Source-grounded answer cards.
|
| 74 |
+
- Confidence label and latency.
|
| 75 |
+
- Retrieved chunk/source metadata.
|
| 76 |
+
- Copy/export actions and read-aloud audio.
|
| 77 |
+
- Agent trace and audit-style context where agentic modes are used.
|
| 78 |
|
| 79 |
+
## Evaluation Notes
|
| 80 |
|
| 81 |
+
The runtime includes:
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 82 |
|
| 83 |
+
- 120 domain prompts in `evaluation/evaluation_queries.md`.
|
| 84 |
+
- 120 multilingual prompts in `evaluation/evaluation_multilingual.md`.
|
| 85 |
+
- Runner and summarizer scripts for repeatable evaluation.
|
| 86 |
+
- Committed snapshots under `evaluation/results`.
|
| 87 |
+
- A generated report under `evaluation/reports/latest_portfolio_report.md`.
|
| 88 |
+
- Decision-critical tests under `tests/`.
|
| 89 |
|
| 90 |
+
Latest committed snapshots show about 2.03s average latency for available rows in both domain and multilingual evaluation exports.
|
|
|
|
|
|
|
| 91 |
|
| 92 |
+
## Limitations
|
| 93 |
|
| 94 |
+
- Current retrieval is FAISS dense vector retrieval. BM25 and reciprocal rank fusion are roadmap work unless implemented later.
|
| 95 |
+
- Fine-Tuned mode requires a configured endpoint or compatible local runtime.
|
| 96 |
+
- Streamlit session state is suitable for live demos but not durable production memory by itself.
|
| 97 |
+
- Outputs are educational and must not be treated as legal, investment, financial, or compliance advice.
|
02-qa-dataset/README.md
CHANGED
|
@@ -1,69 +1,87 @@
|
|
| 1 |
-
#
|
| 2 |
|
| 3 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 4 |
|
| 5 |
## Dataset
|
| 6 |
-
[banking-finance-qa-dataset](https://huggingface.co/datasets/RakeshMadasani/banking-finance-qa-dataset)
|
| 7 |
|
| 8 |
-
|
| 9 |
-
Published dataset page with train and validation splits:
|
| 10 |
|
| 11 |

|
| 12 |
|
| 13 |
## Summary
|
| 14 |
-
|
| 15 |
-
|
| 16 |
-
-
|
| 17 |
-
|
| 18 |
-
|
| 19 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 20 |
|
| 21 |
## Coverage
|
| 22 |
-
|
| 23 |
-
- CDD
|
| 24 |
-
- FDIC
|
| 25 |
-
- Basel III
|
| 26 |
-
- RBI
|
| 27 |
-
- SAR
|
| 28 |
-
-
|
| 29 |
-
|
| 30 |
-
|
| 31 |
-
|
| 32 |
-
|
| 33 |
-
-
|
| 34 |
-
-
|
| 35 |
-
|
| 36 |
-
|
| 37 |
-
|
| 38 |
-
|
| 39 |
-
|
| 40 |
-
|
| 41 |
-
|
| 42 |
-
|
| 43 |
-
|
| 44 |
-
|
| 45 |
-
|
| 46 |
-
|
| 47 |
-
## Example Schema
|
| 48 |
-
|
| 49 |
-
```json
|
| 50 |
-
{
|
| 51 |
-
"instruction": "What is the FDIC deposit insurance limit in the United States?",
|
| 52 |
-
"input": "",
|
| 53 |
-
"output": "The FDIC insures deposits up to $250,000 per depositor, per insured bank, per account ownership category."
|
| 54 |
-
}
|
| 55 |
```
|
| 56 |
|
| 57 |
-
|
| 58 |
-
|
| 59 |
-
|
| 60 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 61 |
|
| 62 |
-
|
| 63 |
-
This project turns raw banking and compliance material into a reusable ML asset rather than stopping at prompt experimentation. It shows the data layer behind the model and application work in the rest of the portfolio.
|
| 64 |
|
| 65 |
## Limitations
|
| 66 |
|
| 67 |
-
-
|
| 68 |
-
-
|
| 69 |
-
-
|
|
|
|
| 1 |
+
# BankingQA Dataset
|
| 2 |
|
| 3 |
+
This folder contains the BankingQA-3K instruction dataset workflow used by the Banking & Finance AI Agent.
|
| 4 |
+
|
| 5 |
+
## Purpose
|
| 6 |
+
|
| 7 |
+
Create a reusable domain dataset for banking and financial compliance instruction tuning. The dataset turns curated banking material into structured instruction-response pairs that support the downstream QLoRA adaptation workflow.
|
| 8 |
|
| 9 |
## Dataset
|
|
|
|
| 10 |
|
| 11 |
+
[banking-finance-qa-dataset](https://huggingface.co/datasets/RakeshMadasani/banking-finance-qa-dataset)
|
|
|
|
| 12 |
|
| 13 |

|
| 14 |
|
| 15 |
## Summary
|
| 16 |
+
|
| 17 |
+
| Item | Value |
|
| 18 |
+
|---|---:|
|
| 19 |
+
| Total examples | 3,002 |
|
| 20 |
+
| Train samples | 2,701 |
|
| 21 |
+
| Validation samples | 301 |
|
| 22 |
+
| Format | Alpaca-style instruction data |
|
| 23 |
+
| Language | English |
|
| 24 |
+
| Domain | Banking, finance, AML, KYC, compliance |
|
| 25 |
+
|
| 26 |
+
## Architecture
|
| 27 |
+
|
| 28 |
+
```mermaid
|
| 29 |
+
flowchart LR
|
| 30 |
+
A["Curated banking material"] --> B["Question generation"]
|
| 31 |
+
B --> C["Instruction / input / output schema"]
|
| 32 |
+
C --> D["Validation and duplicate checks"]
|
| 33 |
+
D --> E["Train / validation split"]
|
| 34 |
+
E --> F["Hugging Face Dataset"]
|
| 35 |
+
```
|
| 36 |
|
| 37 |
## Coverage
|
| 38 |
+
|
| 39 |
+
- AML, KYC, CDD, and EDD.
|
| 40 |
+
- FDIC deposit insurance.
|
| 41 |
+
- Basel III capital concepts.
|
| 42 |
+
- RBI and India banking compliance topics.
|
| 43 |
+
- SAR, CTR, transaction monitoring, and financial crime concepts.
|
| 44 |
+
- General banking and finance fundamentals.
|
| 45 |
+
|
| 46 |
+
## Key Files
|
| 47 |
+
|
| 48 |
+
| File | Role |
|
| 49 |
+
|---|---|
|
| 50 |
+
| `generate_dataset.py` | Builds instruction-response examples from curated material. |
|
| 51 |
+
| `validate_dataset.py` | Checks structure, duplicates, and dataset quality signals. |
|
| 52 |
+
| `upload_to_hf.py` | Publishes dataset artifacts and dataset card to Hugging Face. |
|
| 53 |
+
| `screenshots/` | Published dataset page screenshots for portfolio review. |
|
| 54 |
+
|
| 55 |
+
## How To Run
|
| 56 |
+
|
| 57 |
+
```bash
|
| 58 |
+
cd 02-qa-dataset
|
| 59 |
+
python generate_dataset.py
|
| 60 |
+
python validate_dataset.py
|
| 61 |
+
python upload_to_hf.py
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 62 |
```
|
| 63 |
|
| 64 |
+
Set `HF_TOKEN` before upload if publishing to the Hub.
|
| 65 |
+
|
| 66 |
+
## Inputs And Outputs
|
| 67 |
+
|
| 68 |
+
Inputs:
|
| 69 |
+
|
| 70 |
+
- Curated banking and compliance source material.
|
| 71 |
+
- Topic coverage plan across AML, KYC, FDIC, RBI, Basel III, and banking operations.
|
| 72 |
+
|
| 73 |
+
Outputs:
|
| 74 |
+
|
| 75 |
+
- Alpaca-style dataset with `instruction`, `input`, and `output` fields.
|
| 76 |
+
- Train and validation splits.
|
| 77 |
+
- Hugging Face dataset repository.
|
| 78 |
+
|
| 79 |
+
## Evaluation Notes
|
| 80 |
|
| 81 |
+
The dataset is validated structurally before publishing. It is designed for fine-tuning and domain adaptation, not as a formal legal or regulatory authority. Downstream quality should be measured through model evaluation and grounded answer testing.
|
|
|
|
| 82 |
|
| 83 |
## Limitations
|
| 84 |
|
| 85 |
+
- The dataset is English-only in its current published form.
|
| 86 |
+
- Coverage is strongest for banking/compliance concepts represented in the curated material.
|
| 87 |
+
- Dataset answers should be treated as training material, not official regulatory advice.
|
03-qlora-finetuning/README.md
CHANGED
|
@@ -1,79 +1,83 @@
|
|
| 1 |
-
#
|
| 2 |
|
| 3 |
-
This
|
|
|
|
|
|
|
|
|
|
|
|
|
| 4 |
|
| 5 |
## Model
|
|
|
|
| 6 |
[banking-finance-mistral-qlora](https://huggingface.co/RakeshMadasani/banking-finance-mistral-qlora)
|
| 7 |
|
| 8 |
-
## Published Model Page
|
| 9 |

|
| 10 |
|
| 11 |
-
##
|
| 12 |
|
| 13 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 14 |
|
| 15 |
-
-
|
| 16 |
-
- `What are the three stages of money laundering?`
|
| 17 |
-
- `What is the difference between AML and KYC?`
|
| 18 |
-
|
| 19 |
-
These prompts are short, easy to judge, and representative of the banking/compliance domain adaptation shown by the model.
|
| 20 |
|
| 21 |
-
|
| 22 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 23 |
|
| 24 |
-
##
|
| 25 |
-
- Method: QLoRA
|
| 26 |
-
- Quantization: 4-bit NF4
|
| 27 |
-
- LoRA rank: 16
|
| 28 |
-
- LoRA alpha: 32
|
| 29 |
-
- LoRA dropout: 0.05
|
| 30 |
-
- Training samples: 2,701
|
| 31 |
-
- Validation samples: 301
|
| 32 |
-
- Global steps: 676
|
| 33 |
-
- Final train loss: 1.13
|
| 34 |
|
| 35 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 36 |
|
| 37 |
-
|
| 38 |
-
- designed for practical notebook GPU usage rather than full distributed training
|
| 39 |
-
- built around an adapter approach so a 7B model could be adapted without full fine-tuning
|
| 40 |
|
| 41 |
-
|
| 42 |
-
- `Banking_QLoRA_Mistral7B_updated.ipynb` - main notebook used for the QLoRA workflow
|
| 43 |
-
- `inference_demo.py` - lightweight inference script for loading the published adapter and testing example prompts
|
| 44 |
|
| 45 |
-
|
|
|
|
|
|
|
|
|
|
| 46 |
|
| 47 |
-
|
| 48 |
-
- **LoRA adapters:** updates a small trainable parameter set instead of full-model weights
|
| 49 |
-
- **Cost-efficient experimentation:** a good fit for domain adaptation when full fine-tuning is too heavy
|
| 50 |
|
| 51 |
-
##
|
| 52 |
-
- parameter-efficient fine-tuning
|
| 53 |
-
- PEFT/LoRA configuration
|
| 54 |
-
- domain adaptation using custom data
|
| 55 |
-
- Hugging Face model publishing
|
| 56 |
|
| 57 |
-
|
| 58 |
|
| 59 |
-
|
| 60 |
-
-
|
| 61 |
-
-
|
| 62 |
-
- Basel-related banking questions
|
| 63 |
|
| 64 |
-
|
| 65 |
|
| 66 |
-
|
|
|
|
|
|
|
| 67 |
|
| 68 |
-
|
| 69 |
-
|---|---|---|
|
| 70 |
-
| Banking definitions | usually reasonable but generic | more direct banking-specific phrasing |
|
| 71 |
-
| Compliance language | broad but sometimes high-level | more targeted AML / KYC / SAR / CTR terminology |
|
| 72 |
-
| India-focused regulation | can be vague or mix jurisdictions | better alignment with RBI-oriented wording from the custom dataset |
|
| 73 |
|
| 74 |
-
|
| 75 |
-
This project shows that the portfolio goes beyond app development and dataset creation into actual model adaptation. Publishing the adapter with a model card, tokenizer files, and LoRA configuration makes the work visible and inspectable as a real model artifact.
|
| 76 |
|
| 77 |
-
##
|
| 78 |
|
| 79 |
-
This
|
|
|
|
|
|
|
|
|
| 1 |
+
# Domain Model Adaptation
|
| 2 |
|
| 3 |
+
This folder contains the QLoRA fine-tuning workflow used to adapt Mistral-7B-Instruct-v0.3 to banking and financial compliance terminology.
|
| 4 |
+
|
| 5 |
+
## Purpose
|
| 6 |
+
|
| 7 |
+
Show the model adaptation layer behind the Banking & Finance AI Agent. The goal is not only to call external APIs, but to demonstrate data preparation, parameter-efficient fine-tuning, adapter publishing, and inference testing.
|
| 8 |
|
| 9 |
## Model
|
| 10 |
+
|
| 11 |
[banking-finance-mistral-qlora](https://huggingface.co/RakeshMadasani/banking-finance-mistral-qlora)
|
| 12 |
|
|
|
|
| 13 |

|
| 14 |
|
| 15 |
+
## Architecture
|
| 16 |
|
| 17 |
+
```mermaid
|
| 18 |
+
flowchart LR
|
| 19 |
+
A["BankingQA-3K dataset"] --> B["Prompt formatting"]
|
| 20 |
+
B --> C["Mistral-7B-Instruct-v0.3"]
|
| 21 |
+
C --> D["4-bit NF4 quantization"]
|
| 22 |
+
D --> E["LoRA adapter training"]
|
| 23 |
+
E --> F["Validation / training metrics"]
|
| 24 |
+
F --> G["Published Hugging Face adapter"]
|
| 25 |
+
```
|
| 26 |
|
| 27 |
+
## Fine-Tuning Summary
|
|
|
|
|
|
|
|
|
|
|
|
|
| 28 |
|
| 29 |
+
| Item | Value |
|
| 30 |
+
|---|---|
|
| 31 |
+
| Base model | `mistralai/Mistral-7B-Instruct-v0.3` |
|
| 32 |
+
| Method | QLoRA |
|
| 33 |
+
| Quantization | 4-bit NF4 |
|
| 34 |
+
| LoRA rank | 16 |
|
| 35 |
+
| LoRA alpha | 32 |
|
| 36 |
+
| LoRA dropout | 0.05 |
|
| 37 |
+
| Training samples | 2,701 |
|
| 38 |
+
| Validation samples | 301 |
|
| 39 |
+
| Global steps | 676 |
|
| 40 |
+
| Final train loss | 1.13 |
|
| 41 |
|
| 42 |
+
## Key Files
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 43 |
|
| 44 |
+
| File | Role |
|
| 45 |
+
|---|---|
|
| 46 |
+
| `Banking_QLoRA_Mistral7B_updated.ipynb` | Main notebook for dataset formatting, QLoRA setup, training, and publishing. |
|
| 47 |
+
| `inference_demo.py` | Lightweight script for loading the adapter and testing domain prompts. |
|
| 48 |
+
| `screenshots/` | Published model page and training-progress screenshots. |
|
| 49 |
|
| 50 |
+
## How To Run
|
|
|
|
|
|
|
| 51 |
|
| 52 |
+
The notebook is the main training artifact. For local inference, use:
|
|
|
|
|
|
|
| 53 |
|
| 54 |
+
```bash
|
| 55 |
+
cd 03-qlora-finetuning
|
| 56 |
+
python inference_demo.py
|
| 57 |
+
```
|
| 58 |
|
| 59 |
+
You need access to the base model, the published adapter, and a compatible local GPU/CPU environment. Full 7B inference can be heavy on consumer machines.
|
|
|
|
|
|
|
| 60 |
|
| 61 |
+
## Inputs And Outputs
|
|
|
|
|
|
|
|
|
|
|
|
|
| 62 |
|
| 63 |
+
Inputs:
|
| 64 |
|
| 65 |
+
- BankingQA-3K instruction dataset.
|
| 66 |
+
- Mistral-7B-Instruct-v0.3 base model.
|
| 67 |
+
- PEFT/QLoRA configuration.
|
|
|
|
| 68 |
|
| 69 |
+
Outputs:
|
| 70 |
|
| 71 |
+
- Published LoRA adapter.
|
| 72 |
+
- Training metrics.
|
| 73 |
+
- Model card and inference demo path.
|
| 74 |
|
| 75 |
+
## Evaluation Notes
|
|
|
|
|
|
|
|
|
|
|
|
|
| 76 |
|
| 77 |
+
The current folder documents training configuration and final train loss. A stronger future benchmark should compare the base model, adapter model, OpenAI path, and Auto routing on the same held-out banking evaluation pack.
|
|
|
|
| 78 |
|
| 79 |
+
## Limitations
|
| 80 |
|
| 81 |
+
- This publishes an adapter artifact, not a fully hosted standalone inference service.
|
| 82 |
+
- Runtime quality depends on loading the compatible base model plus adapter correctly.
|
| 83 |
+
- The model should be evaluated on held-out prompts before making strong accuracy claims.
|
04-conversational-memory/README.md
CHANGED
|
@@ -1,139 +1,51 @@
|
|
| 1 |
-
#
|
| 2 |
|
| 3 |
-
This
|
| 4 |
|
| 5 |
-
##
|
| 6 |
|
| 7 |
-
-
|
| 8 |
-
- maintains per-session conversation history for multi-turn follow-up questions
|
| 9 |
-
- truncates and summarizes long sessions to control prompt growth
|
| 10 |
-
- exposes clean API endpoints for chat, session clearing, and health checks
|
| 11 |
-
- supports switchable LLM backends via environment variable
|
| 12 |
-
- supports side-by-side backend comparison on the same retrieved context
|
| 13 |
-
- supports evaluation of memory-on vs memory-off coherence
|
| 14 |
|
| 15 |
## Architecture
|
| 16 |
|
| 17 |
```mermaid
|
| 18 |
flowchart LR
|
| 19 |
-
A["Client
|
| 20 |
-
A -->
|
| 21 |
-
B -->
|
| 22 |
-
|
| 23 |
-
|
| 24 |
-
|
| 25 |
-
|
| 26 |
-
F -->
|
| 27 |
-
|
| 28 |
-
|
| 29 |
-
|
| 30 |
-
|
| 31 |
-
I --> L["HF answer"]
|
| 32 |
-
K --> M["Side-by-side comparison output"]
|
| 33 |
-
L --> M
|
| 34 |
-
|
| 35 |
-
N["DELETE /session/{session_id}"] --> C
|
| 36 |
-
O["GET /health"] --> B
|
| 37 |
```
|
| 38 |
|
| 39 |
-
##
|
| 40 |
-
|
| 41 |
-
```text
|
| 42 |
-
04-conversational-memory/
|
| 43 |
-
|-- app/
|
| 44 |
-
| |-- __init__.py
|
| 45 |
-
| |-- main.py
|
| 46 |
-
| |-- memory.py
|
| 47 |
-
| |-- models.py
|
| 48 |
-
| |-- rag_chain.py
|
| 49 |
-
| `-- summarizer.py
|
| 50 |
-
|-- evaluation/
|
| 51 |
-
| |-- coherence_eval.py
|
| 52 |
-
| `-- questions.csv
|
| 53 |
-
|-- tests/
|
| 54 |
-
| `-- test_memory.py
|
| 55 |
-
|-- .env.example
|
| 56 |
-
|-- README.md
|
| 57 |
-
`-- requirements.txt
|
| 58 |
-
```
|
| 59 |
-
|
| 60 |
-
## API Endpoints
|
| 61 |
-
|
| 62 |
-
### `POST /chat`
|
| 63 |
-
Accepts a user message plus an optional `session_id` and returns:
|
| 64 |
-
|
| 65 |
-
- `session_id`
|
| 66 |
-
- `response`
|
| 67 |
-
- `turn_count`
|
| 68 |
-
- `sources`
|
| 69 |
-
- `confidence`
|
| 70 |
-
- `history_used`
|
| 71 |
-
- `summary_used`
|
| 72 |
-
|
| 73 |
-
### `DELETE /session/{session_id}`
|
| 74 |
-
Clears the stored memory for a given session.
|
| 75 |
-
|
| 76 |
-
### `GET /health`
|
| 77 |
-
Returns a lightweight health status and current session count.
|
| 78 |
-
|
| 79 |
-
### `POST /chat/compare`
|
| 80 |
-
Runs both backends on the same question and retrieved context, then returns both answers side by side for direct comparison.
|
| 81 |
|
| 82 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 83 |
|
| 84 |
-
##
|
| 85 |
-
|
| 86 |
-
- Python 3.10+
|
| 87 |
-
- banking knowledge files available in `01-rag-system/`
|
| 88 |
-
- one backend configured:
|
| 89 |
-
- `openai` with `OPENAI_API_KEY`
|
| 90 |
-
- `local_hf` with model + adapter access and enough local GPU/compute
|
| 91 |
-
|
| 92 |
-
### Install
|
| 93 |
|
| 94 |
```bash
|
|
|
|
| 95 |
pip install -r requirements.txt
|
| 96 |
-
|
| 97 |
-
|
| 98 |
-
### Environment
|
| 99 |
-
|
| 100 |
-
```bash
|
| 101 |
-
LLM_BACKEND=openai
|
| 102 |
-
OPENAI_API_KEY=your_openai_api_key_here
|
| 103 |
-
OPENAI_MODEL=gpt-4o-mini
|
| 104 |
-
LOCAL_HF_MODEL_ID=mistralai/Mistral-7B-Instruct-v0.3
|
| 105 |
-
LOCAL_HF_ADAPTER_ID=RakeshMadasani/banking-finance-mistral-qlora
|
| 106 |
-
LOCAL_HF_DEVICE=auto
|
| 107 |
-
```
|
| 108 |
-
|
| 109 |
-
### Backend options
|
| 110 |
-
|
| 111 |
-
#### Option A - OpenAI backend
|
| 112 |
-
|
| 113 |
-
```bash
|
| 114 |
-
LLM_BACKEND=openai
|
| 115 |
-
OPENAI_API_KEY=your_openai_api_key_here
|
| 116 |
-
OPENAI_MODEL=gpt-4o-mini
|
| 117 |
-
```
|
| 118 |
-
|
| 119 |
-
#### Option B - Local HF backend
|
| 120 |
-
|
| 121 |
-
```bash
|
| 122 |
-
LLM_BACKEND=local_hf
|
| 123 |
-
LOCAL_HF_MODEL_ID=mistralai/Mistral-7B-Instruct-v0.3
|
| 124 |
-
LOCAL_HF_ADAPTER_ID=RakeshMadasani/banking-finance-mistral-qlora
|
| 125 |
-
LOCAL_HF_DEVICE=auto
|
| 126 |
-
```
|
| 127 |
-
|
| 128 |
-
This path is stronger for the portfolio because it links Project 3 directly into Project 4, but it also requires a compatible local environment and enough compute to load the base model plus adapter.
|
| 129 |
-
|
| 130 |
-
### Run the API
|
| 131 |
-
|
| 132 |
-
```bash
|
| 133 |
uvicorn app.main:app --reload --port 8000
|
| 134 |
```
|
| 135 |
|
| 136 |
-
|
| 137 |
|
| 138 |
```bash
|
| 139 |
curl -X POST "http://127.0.0.1:8000/chat" ^
|
|
@@ -141,88 +53,29 @@ curl -X POST "http://127.0.0.1:8000/chat" ^
|
|
| 141 |
-d "{\"message\":\"What is KYC?\",\"session_id\":\"demo-session\",\"use_memory\":true}"
|
| 142 |
```
|
| 143 |
|
| 144 |
-
##
|
| 145 |
-
|
| 146 |
-
```bash
|
| 147 |
-
curl -X POST "http://127.0.0.1:8000/chat/compare" ^
|
| 148 |
-
-H "Content-Type: application/json" ^
|
| 149 |
-
-d "{\"message\":\"What is the difference between AML and KYC?\",\"session_id\":\"compare-session\",\"use_memory\":true}"
|
| 150 |
-
```
|
| 151 |
-
|
| 152 |
-
## Evaluation
|
| 153 |
-
|
| 154 |
-
The `evaluation/coherence_eval.py` script compares memory-off and memory-on responses on multi-turn conversations such as:
|
| 155 |
-
|
| 156 |
-
- KYC followed by follow-up questions about its components and its relation to AML
|
| 157 |
-
- Basel III followed by capital requirement and Basel II comparison prompts
|
| 158 |
-
- SAR followed by reporting-threshold clarification
|
| 159 |
-
|
| 160 |
-
Run:
|
| 161 |
-
|
| 162 |
-
```bash
|
| 163 |
-
python evaluation/coherence_eval.py
|
| 164 |
-
```
|
| 165 |
-
|
| 166 |
-
This prints:
|
| 167 |
-
|
| 168 |
-
- average coherence score with memory off
|
| 169 |
-
- average coherence score with memory on
|
| 170 |
-
- relative improvement
|
| 171 |
-
|
| 172 |
-
Use the resulting percentage as evidence for a claim like `17% improvement` only when your real run produces that number.
|
| 173 |
-
|
| 174 |
-
## Execution Proof
|
| 175 |
-
|
| 176 |
-
This project has already been exercised locally at a lightweight level:
|
| 177 |
-
|
| 178 |
-
- `/health` returned a successful status response
|
| 179 |
-
- `/chat` returned a real answer
|
| 180 |
-
- a follow-up `/chat` call on the same session showed `history_used = true`
|
| 181 |
-
- `/chat/compare` is implemented and reachable, with the local HF runtime path dependent on the local model environment
|
| 182 |
-
|
| 183 |
-
For the current execution record, see [`evaluation/results.md`](./evaluation/results.md).
|
| 184 |
-
|
| 185 |
-
## Design Decisions
|
| 186 |
-
|
| 187 |
-
### Why truncation + summarization instead of only windowing?
|
| 188 |
-
|
| 189 |
-
- simple windowing drops older context entirely
|
| 190 |
-
- summarization preserves earlier intent and constraints
|
| 191 |
-
- retaining the most recent turns keeps the API responsive for active follow-ups
|
| 192 |
-
- current session store is in-process for demo simplicity; a production upgrade path would be Redis or Postgres
|
| 193 |
-
|
| 194 |
-
### Why reuse the Project 1 knowledge base?
|
| 195 |
-
|
| 196 |
-
- keeps the backend directly tied to your deployed banking RAG system
|
| 197 |
-
- strengthens the portfolio story across app layer and backend layer
|
| 198 |
-
- makes evaluation easier because the data and retrieval domain stay consistent
|
| 199 |
-
|
| 200 |
-
### Why support switchable backends?
|
| 201 |
-
|
| 202 |
-
- OpenAI is the fastest path to ship and demo the architecture
|
| 203 |
-
- a local HF backend creates a stronger data -> model -> backend portfolio chain
|
| 204 |
-
- the same FastAPI layer can serve both paths without changing the API contract
|
| 205 |
|
| 206 |
-
|
| 207 |
|
| 208 |
-
-
|
| 209 |
-
-
|
| 210 |
-
-
|
|
|
|
| 211 |
|
| 212 |
-
|
| 213 |
|
| 214 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 215 |
|
| 216 |
-
|
| 217 |
-
- Project 2 provides the banking instruction dataset
|
| 218 |
-
- Project 3 provides the fine-tuned model artifact
|
| 219 |
-
- Project 4 adds the backend memory and orchestration layer for multi-turn interaction
|
| 220 |
|
| 221 |
-
|
| 222 |
|
| 223 |
-
|
| 224 |
|
| 225 |
-
-
|
| 226 |
-
-
|
| 227 |
-
-
|
| 228 |
-
- the current state of `/chat/compare` validation
|
|
|
|
| 1 |
+
# Memory And Orchestration Backend
|
| 2 |
|
| 3 |
+
This folder contains a FastAPI conversational memory backend for the Banking & Finance AI Agent. It adds session handling, controlled history retention, summarization, and backend comparison around the same banking knowledge workflow.
|
| 4 |
|
| 5 |
+
## Purpose
|
| 6 |
|
| 7 |
+
Move the system beyond stateless single-turn answers by introducing reusable API endpoints and memory-aware orchestration. This layer is separate from the live Streamlit Space so it can evolve toward production API deployment without destabilizing the public demo.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 8 |
|
| 9 |
## Architecture
|
| 10 |
|
| 11 |
```mermaid
|
| 12 |
flowchart LR
|
| 13 |
+
A["Client or frontend"] --> B["POST /chat"]
|
| 14 |
+
A --> C["POST /chat/compare"]
|
| 15 |
+
B --> D["Session memory store"]
|
| 16 |
+
C --> D
|
| 17 |
+
D --> E["History truncation / summarization"]
|
| 18 |
+
E --> F["Shared banking retrieval"]
|
| 19 |
+
F --> G["OpenAI backend"]
|
| 20 |
+
F --> H["Local HF adapter backend"]
|
| 21 |
+
G --> I["Response"]
|
| 22 |
+
H --> J["Comparison response"]
|
| 23 |
+
K["DELETE /session/{id}"] --> D
|
| 24 |
+
L["GET /health"] --> B
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 25 |
```
|
| 26 |
|
| 27 |
+
## Key Files
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 28 |
|
| 29 |
+
| File | Role |
|
| 30 |
+
|---|---|
|
| 31 |
+
| `app/main.py` | FastAPI application and route definitions. |
|
| 32 |
+
| `app/memory.py` | Session memory store and state handling. |
|
| 33 |
+
| `app/rag_chain.py` | Retrieval and backend generation logic. |
|
| 34 |
+
| `app/summarizer.py` | Conversation summarization support. |
|
| 35 |
+
| `app/models.py` | Request and response models. |
|
| 36 |
+
| `evaluation/coherence_eval.py` | Memory-on versus memory-off coherence evaluation. |
|
| 37 |
+
| `tests/test_memory.py` | Unit tests for memory behavior. |
|
| 38 |
|
| 39 |
+
## How To Run
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 40 |
|
| 41 |
```bash
|
| 42 |
+
cd 04-conversational-memory
|
| 43 |
pip install -r requirements.txt
|
| 44 |
+
copy .env.example .env
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 45 |
uvicorn app.main:app --reload --port 8000
|
| 46 |
```
|
| 47 |
|
| 48 |
+
Example:
|
| 49 |
|
| 50 |
```bash
|
| 51 |
curl -X POST "http://127.0.0.1:8000/chat" ^
|
|
|
|
| 53 |
-d "{\"message\":\"What is KYC?\",\"session_id\":\"demo-session\",\"use_memory\":true}"
|
| 54 |
```
|
| 55 |
|
| 56 |
+
## Inputs And Outputs
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 57 |
|
| 58 |
+
Inputs:
|
| 59 |
|
| 60 |
+
- User message.
|
| 61 |
+
- Optional `session_id`.
|
| 62 |
+
- Optional memory usage flag.
|
| 63 |
+
- Backend configuration via environment variables.
|
| 64 |
|
| 65 |
+
Outputs:
|
| 66 |
|
| 67 |
+
- Session-aware answer.
|
| 68 |
+
- Sources and confidence.
|
| 69 |
+
- Turn count.
|
| 70 |
+
- Flags showing whether history or summary was used.
|
| 71 |
+
- Side-by-side backend comparison through `/chat/compare`.
|
| 72 |
|
| 73 |
+
## Evaluation Notes
|
|
|
|
|
|
|
|
|
|
| 74 |
|
| 75 |
+
The coherence evaluation compares memory-off and memory-on behavior across follow-up conversations. Claims such as percentage improvement should only be made from actual generated evaluation output, not assumed values.
|
| 76 |
|
| 77 |
+
## Limitations
|
| 78 |
|
| 79 |
+
- The current memory store is lightweight and in-process; production use should move to Redis, Postgres, or another durable store.
|
| 80 |
+
- Local Hugging Face adapter inference depends on a compatible environment and sufficient compute.
|
| 81 |
+
- This backend is a system-design layer and is not the live public Space entry point.
|
|
|
CONTRIBUTING.md
ADDED
|
@@ -0,0 +1,62 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Contributing
|
| 2 |
+
|
| 3 |
+
Thanks for improving the Banking & Finance AI Agent. This repository is maintained as a serious AI engineering portfolio, so changes should be small, testable, and honest about what the code supports.
|
| 4 |
+
|
| 5 |
+
## Setup
|
| 6 |
+
|
| 7 |
+
```bash
|
| 8 |
+
git clone https://github.com/rakeshmadasaniai/banking-genai-portfolio.git
|
| 9 |
+
cd banking-genai-portfolio
|
| 10 |
+
python -m venv .venv
|
| 11 |
+
.venv\Scripts\Activate.ps1
|
| 12 |
+
pip install -r 01-rag-system/requirements.txt
|
| 13 |
+
```
|
| 14 |
+
|
| 15 |
+
Copy the sample environment file and fill local secrets:
|
| 16 |
+
|
| 17 |
+
```bash
|
| 18 |
+
copy .env.example .env
|
| 19 |
+
```
|
| 20 |
+
|
| 21 |
+
Never commit real API keys or tokens.
|
| 22 |
+
|
| 23 |
+
## Branching
|
| 24 |
+
|
| 25 |
+
- Use short descriptive branches.
|
| 26 |
+
- Prefer `feature/...`, `fix/...`, or `docs/...`.
|
| 27 |
+
- Keep deployment-risky changes separate from documentation-only changes.
|
| 28 |
+
|
| 29 |
+
## Testing
|
| 30 |
+
|
| 31 |
+
Run the available unit tests:
|
| 32 |
+
|
| 33 |
+
```bash
|
| 34 |
+
cd 01-rag-system
|
| 35 |
+
python -m unittest discover -s tests -p "test_*.py"
|
| 36 |
+
```
|
| 37 |
+
|
| 38 |
+
When changing evaluation logic, also run the relevant scripts under `01-rag-system/evaluation`.
|
| 39 |
+
|
| 40 |
+
## Pull Request Expectations
|
| 41 |
+
|
| 42 |
+
Each PR should include:
|
| 43 |
+
|
| 44 |
+
- What changed.
|
| 45 |
+
- Why it changed.
|
| 46 |
+
- How it was tested.
|
| 47 |
+
- Any limitations or follow-up work.
|
| 48 |
+
|
| 49 |
+
For AI behavior changes, include at least one before/after example and avoid unsupported accuracy claims.
|
| 50 |
+
|
| 51 |
+
## Code Style
|
| 52 |
+
|
| 53 |
+
- Keep Python readable and typed where practical.
|
| 54 |
+
- Prefer small functions over large hidden control flow.
|
| 55 |
+
- Add comments only where the logic is non-obvious.
|
| 56 |
+
- Preserve existing UI design language unless the change is explicitly a redesign.
|
| 57 |
+
|
| 58 |
+
## Documentation Style
|
| 59 |
+
|
| 60 |
+
- Use "AI agent", "grounded GenAI system", or "tool-calling AI system" when accurate.
|
| 61 |
+
- Avoid "AGI" or "fully autonomous" unless the implementation and evaluation clearly support that exact claim.
|
| 62 |
+
- Keep metrics tied to committed evaluation files.
|
EVALUATION.md
ADDED
|
@@ -0,0 +1,77 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Evaluation
|
| 2 |
+
|
| 3 |
+
This document describes how the Banking & Finance AI Agent is evaluated and what should be improved next.
|
| 4 |
+
|
| 5 |
+
## Goals
|
| 6 |
+
|
| 7 |
+
- Measure whether answers are grounded in retrieved banking context.
|
| 8 |
+
- Measure latency for product responsiveness.
|
| 9 |
+
- Compare OpenAI, Fine-Tuned, Auto, and agentic paths where configured.
|
| 10 |
+
- Test multilingual behavior rather than assuming English-only quality.
|
| 11 |
+
- Capture failure modes openly so improvements are traceable.
|
| 12 |
+
|
| 13 |
+
## Evaluation Assets
|
| 14 |
+
|
| 15 |
+
| Asset | Location | Purpose |
|
| 16 |
+
|---|---|---|
|
| 17 |
+
| Domain query pack | `01-rag-system/evaluation/evaluation_queries.md` | Banking, AML, KYC, FDIC, RBI, Basel III, payments, and compliance prompts. |
|
| 18 |
+
| Multilingual query pack | `01-rag-system/evaluation/evaluation_multilingual.md` | Multilingual coverage across model modes. |
|
| 19 |
+
| Runner | `01-rag-system/evaluation/run_eval_sets.py` | Executes evaluation packs. |
|
| 20 |
+
| Summarizer | `01-rag-system/evaluation/summarize_eval_sets.py` | Produces CSV/JSON summaries. |
|
| 21 |
+
| Results | `01-rag-system/evaluation/results/` | Committed result snapshots. |
|
| 22 |
+
| Portfolio report | `01-rag-system/evaluation/reports/latest_portfolio_report.md` | Recruiter-friendly summary generated from committed artifacts. |
|
| 23 |
+
| Autonomy audit | `AUTONOMY_EVALUATION.md` | Honest status of the agentic/autonomous runtime. |
|
| 24 |
+
|
| 25 |
+
## Current Snapshot
|
| 26 |
+
|
| 27 |
+
| Evaluation set | Total prompts | Available evaluated rows | Average latency | Median latency |
|
| 28 |
+
|---|---:|---:|---:|---:|
|
| 29 |
+
| Domain pack | 120 | 80 | 2037.0 ms | 2036.0 ms |
|
| 30 |
+
| Multilingual pack | 120 | 80 | 2031.8 ms | 2031.5 ms |
|
| 31 |
+
|
| 32 |
+
The committed export includes unavailable rows where the relevant backend was not active in the local evaluation environment. This is preserved intentionally so the results stay auditable.
|
| 33 |
+
|
| 34 |
+
## Prompt Categories
|
| 35 |
+
|
| 36 |
+
- Banking definitions and explainers.
|
| 37 |
+
- AML/KYC/CDD/EDD compliance.
|
| 38 |
+
- FDIC deposit insurance.
|
| 39 |
+
- RBI and India banking compliance.
|
| 40 |
+
- Basel III capital and risk concepts.
|
| 41 |
+
- Payments and transaction-monitoring scenarios.
|
| 42 |
+
- Cross-jurisdiction comparison prompts.
|
| 43 |
+
- Multilingual banking questions.
|
| 44 |
+
- Agentic decision scenarios such as fraud, sanctions, short-horizon investing, and life-event planning.
|
| 45 |
+
|
| 46 |
+
## Groundedness Scoring
|
| 47 |
+
|
| 48 |
+
The runtime uses scoring utilities to estimate:
|
| 49 |
+
|
| 50 |
+
- overlap with retrieved documents,
|
| 51 |
+
- completeness of the answer,
|
| 52 |
+
- latency quality,
|
| 53 |
+
- combined candidate score for routing.
|
| 54 |
+
|
| 55 |
+
These scores are useful product signals, not formal legal or regulatory validation.
|
| 56 |
+
|
| 57 |
+
## Latency Measurement
|
| 58 |
+
|
| 59 |
+
Latency is recorded in milliseconds in model results and evaluation outputs. Current committed snapshot averages are around 2.03 seconds for available rows. Future reports should add p50, p95, p99, and per-mode latency.
|
| 60 |
+
|
| 61 |
+
## Failure Modes To Track
|
| 62 |
+
|
| 63 |
+
- Retrieval misses for narrow regulatory facts.
|
| 64 |
+
- Answers that ask for clarification when enough information already exists.
|
| 65 |
+
- Overly generic policy-generation responses.
|
| 66 |
+
- Cross-jurisdiction questions that need explicit table formatting.
|
| 67 |
+
- Voice and upload behavior that depends on browser/runtime permissions.
|
| 68 |
+
- Fine-Tuned mode availability when the hosted endpoint is not configured.
|
| 69 |
+
|
| 70 |
+
## Future Benchmark Plan
|
| 71 |
+
|
| 72 |
+
- Add a gold-answer set for 100 high-value banking and compliance questions.
|
| 73 |
+
- Add multilingual human review for at least five languages.
|
| 74 |
+
- Add separate benchmarks for retrieval-only, OpenAI, Fine-Tuned, Auto, and agentic modes.
|
| 75 |
+
- Add document-upload tests for PDF, DOCX, TXT, and image inputs.
|
| 76 |
+
- Add voice input/output smoke tests where runtime support is available.
|
| 77 |
+
- Track before/after scores for BM25/RRF once sparse retrieval is implemented.
|
README.md
CHANGED
|
@@ -1,5 +1,5 @@
|
|
| 1 |
---
|
| 2 |
-
title:
|
| 3 |
emoji: 🌎
|
| 4 |
colorFrom: blue
|
| 5 |
colorTo: indigo
|
|
@@ -10,290 +10,185 @@ app_file: 01-rag-system/app.py
|
|
| 10 |
pinned: false
|
| 11 |
---
|
| 12 |
|
| 13 |
-
#
|
| 14 |
|
| 15 |
-
|
| 16 |
-
[
|
|
|
|
|
|
|
|
|
|
|
|
|
| 17 |
|
| 18 |
-
|
| 19 |
|
| 20 |
-
I
|
| 21 |
|
| 22 |
-
|
| 23 |
-
- measurable with repeatable evaluation packs
|
| 24 |
-
- flexible across OpenAI, Fine-Tuned, and Auto routing
|
| 25 |
-
- multilingual enough for broader banking users
|
| 26 |
-
- strong enough to discuss as an engineering system, not just a UI
|
| 27 |
|
| 28 |
-
##
|
| 29 |
|
| 30 |
-
|
| 31 |
-
|
| 32 |
-
|
| 33 |
-
|
| 34 |
-
|
| 35 |
-
|
| 36 |
-
|
| 37 |
-
## Live Product
|
| 38 |
-
|
| 39 |
-
- **Live app:** [banking-finance-rag](https://huggingface.co/spaces/RakeshMadasani/banking-finance-rag)
|
| 40 |
-
- **Dataset:** [banking-finance-qa-dataset](https://huggingface.co/datasets/RakeshMadasani/banking-finance-qa-dataset)
|
| 41 |
-
- **Fine-tuned model:** [banking-finance-mistral-qlora](https://huggingface.co/RakeshMadasani/banking-finance-mistral-qlora)
|
| 42 |
-
|
| 43 |
-
## What This Repo Shows
|
| 44 |
-
|
| 45 |
-
| Layer | What is in the repo | Why it matters |
|
| 46 |
-
|---|---|---|
|
| 47 |
-
| Product | A live Banking & Finance Copilot | Shows a shipped, testable AI product |
|
| 48 |
-
| Retrieval | FAISS + banking knowledge + source cards | Keeps answers grounded and explainable |
|
| 49 |
-
| Data | A domain-specific QA dataset | Shows data ownership, not just prompting |
|
| 50 |
-
| Model | QLoRA fine-tuning workflow | Shows model adaptation beyond API usage |
|
| 51 |
-
| Backend | Conversational memory API | Shows system thinking and architecture depth |
|
| 52 |
-
| Evaluation | 120-query domain set + 120-query multilingual set | Shows repeatable measurement, not anecdotal demos |
|
| 53 |
-
|
| 54 |
-
## Where The Product Code Lives
|
| 55 |
-
|
| 56 |
-
If someone lands on this repo from GitHub first, the main product code is not hidden in a separate private service. It lives directly inside [`01-rag-system`](01-rag-system):
|
| 57 |
-
|
| 58 |
-
- [`01-rag-system/core`](01-rag-system/core)
|
| 59 |
-
runtime orchestration, retrieval flow, prompts, and shared utilities
|
| 60 |
-
- [`01-rag-system/features`](01-rag-system/features)
|
| 61 |
-
product UI, uploads, voice output, answer cards, and interaction behavior
|
| 62 |
-
- [`01-rag-system/models`](01-rag-system/models)
|
| 63 |
-
OpenAI mode, Fine-Tuned mode, and Auto routing logic
|
| 64 |
-
|
| 65 |
-
That structure matters because I wanted the repo to read like a real product codebase, not a single README pointing to an external demo.
|
| 66 |
-
|
| 67 |
-
## How The Portfolio Evolves
|
| 68 |
-
|
| 69 |
-
This repo is one system built in layers.
|
| 70 |
-
|
| 71 |
-
### 1. Product layer
|
| 72 |
-
|
| 73 |
-
I started with the user-facing assistant in [`01-rag-system`](01-rag-system). The goal was simple: if someone asks a banking or compliance question, the product should answer clearly and show the evidence behind the answer.
|
| 74 |
-
|
| 75 |
-
### 2. Data layer
|
| 76 |
-
|
| 77 |
-
Once the first retrieval system worked, I created a banking QA dataset in [`02-qa-dataset`](02-qa-dataset) so the domain logic would not live only inside prompts and chunk text.
|
| 78 |
-
|
| 79 |
-
### 3. Model layer
|
| 80 |
-
|
| 81 |
-
Then I fine-tuned a banking-domain adapter in [`03-qlora-finetuning`](03-qlora-finetuning) to show that I can move from application wiring into actual model adaptation.
|
| 82 |
-
|
| 83 |
-
### 4. Backend layer
|
| 84 |
-
|
| 85 |
-
Finally, I added session memory and orchestration work in [`04-conversational-memory`](04-conversational-memory), which made the portfolio feel more like a real product system than a single-page demo.
|
| 86 |
-
|
| 87 |
-
## Architecture
|
| 88 |
|
| 89 |
-
##
|
| 90 |
|
| 91 |
-
|
| 92 |
-
|
| 93 |
-
|
| 94 |
-
|
| 95 |
-
|
| 96 |
-
|
| 97 |
-
|
| 98 |
-
A --> F["Instruction dataset generation"]
|
| 99 |
-
F --> G["Banking QA dataset"]
|
| 100 |
-
G --> H["QLoRA fine-tuning"]
|
| 101 |
-
H --> I["Published banking adapter"]
|
| 102 |
-
|
| 103 |
-
E --> J["OpenAI path"]
|
| 104 |
-
E --> K["Fine-Tuned path"]
|
| 105 |
-
E --> L["Auto routing"]
|
| 106 |
-
|
| 107 |
-
J --> M["Live Banking & Finance Copilot"]
|
| 108 |
-
K --> M
|
| 109 |
-
L --> M
|
| 110 |
-
|
| 111 |
-
M --> N["Sources, confidence, latency, uploads, read aloud"]
|
| 112 |
-
M --> O["Conversational memory backend"]
|
| 113 |
-
```
|
| 114 |
|
| 115 |
-
##
|
| 116 |
|
| 117 |
```mermaid
|
| 118 |
flowchart TD
|
| 119 |
-
U["
|
| 120 |
-
|
| 121 |
-
|
| 122 |
-
|
| 123 |
-
|
| 124 |
-
|
| 125 |
-
|
| 126 |
-
|
| 127 |
-
|
| 128 |
-
|
| 129 |
-
|
| 130 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 131 |
```
|
| 132 |
|
| 133 |
-
##
|
| 134 |
-
|
| 135 |
-
The main product lives in [`01-rag-system`](01-rag-system).
|
| 136 |
-
|
| 137 |
-
What it does well:
|
| 138 |
-
- grounded banking and compliance Q&A
|
| 139 |
-
- OpenAI, Fine-Tuned, and Auto modes
|
| 140 |
-
- source-backed answers with visible evidence
|
| 141 |
-
- PDF, DOCX, TXT, and image upload support
|
| 142 |
-
- multilingual answer support
|
| 143 |
-
- evaluation workflows committed alongside the app
|
| 144 |
-
|
| 145 |
-
### Product Walkthrough
|
| 146 |
-
|
| 147 |
-
#### A clean first impression for the product
|
| 148 |
-
|
| 149 |
-
This is the opening experience of the Banking & Finance Copilot: stable sidebar, mode selector, welcome guidance, and multilingual starter questions that make the product feel usable immediately.
|
| 150 |
-
|
| 151 |
-

|
| 152 |
-
|
| 153 |
-
#### A grounded English answer that feels concise and useful
|
| 154 |
-
|
| 155 |
-
This example shows the assistant answering a KYC question in English with a direct explanation, compact bullets, visible latency, and retrieved source support.
|
| 156 |
|
| 157 |
-
|
| 158 |
|
| 159 |
-
|
| 160 |
-
|
| 161 |
-
|
| 162 |
-
|
| 163 |
-
|
| 164 |
-
|
| 165 |
-
#### Multilingual grounding working in Chinese as well
|
| 166 |
-
|
| 167 |
-
This example shows the same KYC workflow in Chinese, which helps demonstrate that multilingual support is part of the product design, not just a side feature.
|
| 168 |
-
|
| 169 |
-

|
| 170 |
-
|
| 171 |
-
## Evaluation
|
| 172 |
-
|
| 173 |
-
The repo includes two larger evaluation packs inside [`01-rag-system/evaluation`](01-rag-system/evaluation):
|
| 174 |
-
|
| 175 |
-
- `evaluation_queries.md`
|
| 176 |
-
120 domain-specific banking, AML, KYC, Basel III, FDIC, RBI, CECL, and payments questions
|
| 177 |
-
- `evaluation_multilingual.md`
|
| 178 |
-
120 multilingual questions grouped across OpenAI, Fine-Tuned, and Auto modes
|
| 179 |
-
|
| 180 |
-
The folder also includes:
|
| 181 |
-
|
| 182 |
-
- `run_eval_sets.py` to run both evaluation packs automatically
|
| 183 |
-
- `summarize_eval_sets.py` to summarize any generated CSV
|
| 184 |
-
- committed result snapshots in [`01-rag-system/evaluation/results`](01-rag-system/evaluation/results)
|
| 185 |
-
- raw CSV outputs plus JSON summaries, so the runs are inspectable rather than just summarized in prose
|
| 186 |
-
|
| 187 |
-
### Latest committed evaluation snapshots
|
| 188 |
|
| 189 |
-
|
| 190 |
|
| 191 |
-
|
| 192 |
-
|---|---|
|
| 193 |
-
| Total prompts | 120 |
|
| 194 |
-
| Available evaluated rows | 80 |
|
| 195 |
-
| Average latency | 2037.0 ms |
|
| 196 |
-
| Median latency | 2036.0 ms |
|
| 197 |
|
| 198 |
-
|
| 199 |
|
| 200 |
| Metric | Result |
|
| 201 |
|---|---|
|
| 202 |
-
|
|
| 203 |
-
|
|
| 204 |
-
|
|
| 205 |
-
|
|
| 206 |
-
|
| 207 |
-
|
| 208 |
-
|
| 209 |
-
|
|
|
|
| 210 |
|
| 211 |
-
|
| 212 |
-
- the committed run has 80 available rows because the OpenAI path was not active in that local export
|
| 213 |
-
- the Fine-Tuned and Auto paths still completed and produced auditable CSV/JSON artifacts
|
| 214 |
|
| 215 |
-
|
| 216 |
|
| 217 |
-
###
|
| 218 |
|
| 219 |
-
|
| 220 |
|
| 221 |
-
|
| 222 |
-
- Output: `01-rag-system/evaluation/reports/latest_portfolio_report.md`
|
| 223 |
|
| 224 |
-
|
| 225 |
|
| 226 |
-
###
|
| 227 |
|
| 228 |
-
|
| 229 |
|
| 230 |
-
|
| 231 |
|
| 232 |
-
|
| 233 |
|
| 234 |
-
|
| 235 |
-
- `python -m unittest discover -s tests -p "test_*.py"`
|
| 236 |
|
| 237 |
-
|
| 238 |
|
| 239 |
-
|
| 240 |
|
| 241 |
-
|
| 242 |
-
|
| 243 |
-
-
|
| 244 |
-
|
|
|
|
|
|
|
| 245 |
|
| 246 |
-
|
| 247 |
|
| 248 |
-
|
| 249 |
-
|
| 250 |
-
|
| 251 |
-
|
| 252 |
-
|
| 253 |
-
|
| 254 |
-
|
| 255 |
-
|
| 256 |
-
|
| 257 |
-
|
| 258 |
-
|
| 259 |
-
|
| 260 |
-
|
| 261 |
|
| 262 |
-
|
| 263 |
|
| 264 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 265 |
|
| 266 |
-
|
| 267 |
|
| 268 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 269 |
|
| 270 |
-
|
| 271 |
|
| 272 |
-
-
|
| 273 |
-
-
|
| 274 |
-
-
|
| 275 |
-
-
|
| 276 |
-
- `01-rag-system/models/auto_router.py`
|
| 277 |
-
- `01-rag-system/models/openai_mode.py`
|
| 278 |
-
- `01-rag-system/models/finetuned_mode.py`
|
| 279 |
-
- `01-rag-system/evaluation/run_eval_sets.py`
|
| 280 |
-
- `01-rag-system/evaluation/summarize_eval_sets.py`
|
| 281 |
-
- `02-qa-dataset/generate_dataset.py`
|
| 282 |
-
- `03-qlora-finetuning/inference_demo.py`
|
| 283 |
-
- `04-conversational-memory/app/main.py`
|
| 284 |
-
- `04-conversational-memory/app/rag_chain.py`
|
| 285 |
|
| 286 |
-
##
|
| 287 |
|
| 288 |
-
``
|
| 289 |
-
banking-genai-portfolio/
|
| 290 |
-
|-- README.md
|
| 291 |
-
|-- 01-rag-system/
|
| 292 |
-
|-- 02-qa-dataset/
|
| 293 |
-
|-- 03-qlora-finetuning/
|
| 294 |
-
`-- 04-conversational-memory/
|
| 295 |
-
```
|
| 296 |
|
| 297 |
-
##
|
| 298 |
|
| 299 |
-
|
|
|
|
| 1 |
---
|
| 2 |
+
title: Banking & Finance AI Agent
|
| 3 |
emoji: 🌎
|
| 4 |
colorFrom: blue
|
| 5 |
colorTo: indigo
|
|
|
|
| 10 |
pinned: false
|
| 11 |
---
|
| 12 |
|
| 13 |
+
# Banking & Finance AI Agent
|
| 14 |
|
| 15 |
+

|
| 16 |
+

|
| 17 |
+

|
| 18 |
+

|
| 19 |
+

|
| 20 |
+

|
| 21 |
|
| 22 |
+
**Production-grade GenAI system for grounded banking, compliance, and financial knowledge workflows.**
|
| 23 |
|
| 24 |
+
I built this system to answer one question: what does it take to move a GenAI product beyond a chatbot demo and into a reliable, measurable AI system?
|
| 25 |
|
| 26 |
+
The result is a live Banking & Finance AI Agent with retrieval, model routing, domain adaptation, conversational memory work, evaluation packs, multilingual UX, upload workflows, voice support, and an autonomy audit. It is intentionally positioned as an AI agent and grounded GenAI platform, not AGI and not an overclaimed fully autonomous production system.
|
|
|
|
|
|
|
|
|
|
|
|
|
| 27 |
|
| 28 |
+
## Live Links
|
| 29 |
|
| 30 |
+
| Asset | Link |
|
| 31 |
+
|---|---|
|
| 32 |
+
| Live app | [Hugging Face Space](https://huggingface.co/spaces/RakeshMadasani/banking-finance-rag) |
|
| 33 |
+
| GitHub repository | [banking-genai-portfolio](https://github.com/rakeshmadasaniai/banking-genai-portfolio) |
|
| 34 |
+
| Fine-tuned model | [banking-finance-mistral-qlora](https://huggingface.co/RakeshMadasani/banking-finance-mistral-qlora) |
|
| 35 |
+
| Dataset | [banking-finance-qa-dataset](https://huggingface.co/datasets/RakeshMadasani/banking-finance-qa-dataset) |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 36 |
|
| 37 |
+
## What This System Does
|
| 38 |
|
| 39 |
+
- Answers banking, finance, AML, KYC, FDIC, Basel III, RBI, and compliance questions with retrieved context.
|
| 40 |
+
- Supports OpenAI, Fine-Tuned, Auto, Agentic Workspace, and Autonomous Max paths where configured.
|
| 41 |
+
- Renders source-grounded answer cards with latency, confidence, retrieved chunks, source cards, copy/export actions, and read-aloud controls.
|
| 42 |
+
- Accepts text, document uploads, image-supported workflows, multilingual prompts, and voice input/output paths.
|
| 43 |
+
- Includes a published BankingQA-3K dataset and a QLoRA Mistral-7B adapter for domain model adaptation.
|
| 44 |
+
- Ships repeatable evaluation packs, committed result snapshots, and an autonomy evaluation note instead of only screenshots.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 45 |
|
| 46 |
+
## Architecture Overview
|
| 47 |
|
| 48 |
```mermaid
|
| 49 |
flowchart TD
|
| 50 |
+
U["User input: text, document, image, or voice"] --> IR["Input router"]
|
| 51 |
+
IR --> UP["Upload and document parsing"]
|
| 52 |
+
IR --> VI["Voice input path"]
|
| 53 |
+
IR --> Q["Normalized user query"]
|
| 54 |
+
|
| 55 |
+
UP --> KB["Runtime knowledge context"]
|
| 56 |
+
Q --> RC["Retrieval coordinator"]
|
| 57 |
+
KB --> RC
|
| 58 |
+
|
| 59 |
+
RC --> FAISS["Implemented: FAISS dense vector search"]
|
| 60 |
+
RC -. "roadmap" .-> BM25["Planned: BM25 sparse search"]
|
| 61 |
+
BM25 -. "roadmap" .-> RRF["Planned: reciprocal rank fusion"]
|
| 62 |
+
FAISS --> GC["Grounded context"]
|
| 63 |
+
RRF -. "future hybrid context" .-> GC
|
| 64 |
+
|
| 65 |
+
GC --> ORCH["LLM orchestration layer"]
|
| 66 |
+
ORCH --> OAI["OpenAI mode"]
|
| 67 |
+
ORCH --> FT["Fine-Tuned mode"]
|
| 68 |
+
ORCH --> AUTO["Auto routing"]
|
| 69 |
+
ORCH --> AGENT["Agentic / Autonomous modes"]
|
| 70 |
+
|
| 71 |
+
OAI --> EVAL["Evaluation + confidence scoring"]
|
| 72 |
+
FT --> EVAL
|
| 73 |
+
AUTO --> EVAL
|
| 74 |
+
AGENT --> EVAL
|
| 75 |
+
|
| 76 |
+
EVAL --> RESP["Response with sources, confidence, latency, and actions"]
|
| 77 |
+
RESP --> MEM["Session memory and audit context"]
|
| 78 |
+
MEM --> ORCH
|
| 79 |
```
|
| 80 |
|
| 81 |
+
## System Design
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 82 |
|
| 83 |
+
The system is built in four layers:
|
| 84 |
|
| 85 |
+
| Layer | Runtime area | Purpose |
|
| 86 |
+
|---|---|---|
|
| 87 |
+
| AI agent runtime | `01-rag-system` | Live Streamlit product, retrieval, orchestration, source-grounded UI, uploads, voice, and agent paths. |
|
| 88 |
+
| BankingQA dataset | `02-qa-dataset` | 3,002-pair instruction dataset covering banking, compliance, AML, KYC, Basel III, FDIC, RBI, and finance topics. |
|
| 89 |
+
| Domain model adaptation | `03-qlora-finetuning` | QLoRA workflow for adapting Mistral-7B-Instruct-v0.3 to banking and financial compliance terminology. |
|
| 90 |
+
| Memory and orchestration | `04-conversational-memory` | FastAPI memory backend with session handling, history management, summarization, and backend comparison. |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 91 |
|
| 92 |
+
The current live Space keeps the original folder path `01-rag-system` so existing Hugging Face deployment URLs and `app_file` metadata continue to work. Documentation now positions that folder as the AI agent runtime while preserving compatibility.
|
| 93 |
|
| 94 |
+
## Evaluation Metrics
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 95 |
|
| 96 |
+
These numbers are from committed project files and should be read as traceable project evidence, not marketing claims.
|
| 97 |
|
| 98 |
| Metric | Result |
|
| 99 |
|---|---|
|
| 100 |
+
| Evaluation prompts | 240 total prompts across domain and multilingual packs |
|
| 101 |
+
| Latest available evaluated rows | 160 rows across committed domain and multilingual snapshots |
|
| 102 |
+
| Domain pack average latency | 2037.0 ms |
|
| 103 |
+
| Multilingual pack average latency | 2031.8 ms |
|
| 104 |
+
| Dataset size | 3,002 QA pairs |
|
| 105 |
+
| Fine-tuning method | Mistral-7B QLoRA |
|
| 106 |
+
| Final train loss | 1.13 |
|
| 107 |
+
| Current retrieval implementation | FAISS dense vector retrieval |
|
| 108 |
+
| Sparse retrieval / RRF | Roadmap item, not claimed as live implementation |
|
| 109 |
|
| 110 |
+
For the current autonomy positioning, see [`AUTONOMY_EVALUATION.md`](AUTONOMY_EVALUATION.md). For the generated portfolio report, see [`01-rag-system/evaluation/reports/latest_portfolio_report.md`](01-rag-system/evaluation/reports/latest_portfolio_report.md).
|
|
|
|
|
|
|
| 111 |
|
| 112 |
+
## Key Capabilities
|
| 113 |
|
| 114 |
+
### Grounded Banking Answers
|
| 115 |
|
| 116 |
+
The runtime retrieves banking material before generation, then presents answers with source cards, confidence labels, latency, and chunk metadata.
|
| 117 |
|
| 118 |
+
### Model Orchestration
|
|
|
|
| 119 |
|
| 120 |
+
OpenAI mode provides a stable general path, Fine-Tuned mode connects the domain adapter path where hosted inference is configured, and Auto mode scores candidate answers based on groundedness, completeness, and latency.
|
| 121 |
|
| 122 |
+
### Agentic Runtime Work
|
| 123 |
|
| 124 |
+
The repo includes agentic/autonomous runtime work with tool-style execution traces and autonomy evaluation. This is presented honestly as a tool-calling AI system and agentic workflow layer, not as AGI.
|
| 125 |
|
| 126 |
+
### Domain Data and Model Adaptation
|
| 127 |
|
| 128 |
+
The dataset and QLoRA adapter show the system is not only prompt engineering. It includes a reusable data asset and a domain-adapted model artifact.
|
| 129 |
|
| 130 |
+
### Memory and API Layer
|
|
|
|
| 131 |
|
| 132 |
+
The FastAPI memory backend demonstrates session-aware conversation handling, summarization/truncation, health checks, and backend comparison endpoints.
|
| 133 |
|
| 134 |
+
## Demo Workflow
|
| 135 |
|
| 136 |
+
1. Open the [live Space](https://huggingface.co/spaces/RakeshMadasani/banking-finance-rag).
|
| 137 |
+
2. Ask a banking or compliance question such as `What are the main KYC requirements for banks?`.
|
| 138 |
+
3. Switch modes to compare OpenAI, Fine-Tuned, Auto, and agentic paths where configured.
|
| 139 |
+
4. Upload a PDF, DOCX, or TXT document and ask a document-grounded question.
|
| 140 |
+
5. Inspect confidence, source cards, latency, retrieved chunks, and read-aloud output.
|
| 141 |
+
6. Review evaluation artifacts under `01-rag-system/evaluation`.
|
| 142 |
|
| 143 |
+
## Repository Structure
|
| 144 |
|
| 145 |
+
```text
|
| 146 |
+
banking-genai-portfolio/
|
| 147 |
+
|-- README.md
|
| 148 |
+
|-- ROADMAP.md
|
| 149 |
+
|-- EVALUATION.md
|
| 150 |
+
|-- SYSTEM_DESIGN.md
|
| 151 |
+
|-- CONTRIBUTING.md
|
| 152 |
+
|-- AUTONOMY_EVALUATION.md
|
| 153 |
+
|-- 01-rag-system/ # AI agent runtime and live Streamlit app
|
| 154 |
+
|-- 02-qa-dataset/ # BankingQA-3K dataset build/publish workflow
|
| 155 |
+
|-- 03-qlora-finetuning/ # Mistral-7B QLoRA adaptation workflow
|
| 156 |
+
`-- 04-conversational-memory/ # FastAPI memory and orchestration backend
|
| 157 |
+
```
|
| 158 |
|
| 159 |
+
## Run Locally
|
| 160 |
|
| 161 |
+
```bash
|
| 162 |
+
git clone https://github.com/rakeshmadasaniai/banking-genai-portfolio.git
|
| 163 |
+
cd banking-genai-portfolio
|
| 164 |
+
python -m venv .venv
|
| 165 |
+
.venv\Scripts\Activate.ps1
|
| 166 |
+
pip install -r 01-rag-system/requirements.txt
|
| 167 |
+
copy .env.example .env
|
| 168 |
+
streamlit run 01-rag-system/app.py
|
| 169 |
+
```
|
| 170 |
|
| 171 |
+
Minimum environment:
|
| 172 |
|
| 173 |
+
```env
|
| 174 |
+
OPENAI_API_KEY=
|
| 175 |
+
HF_TOKEN=
|
| 176 |
+
MODEL_MODE=OpenAI
|
| 177 |
+
TOP_K=4
|
| 178 |
+
TEMPERATURE=0.2
|
| 179 |
+
```
|
| 180 |
|
| 181 |
+
## Known Limitations
|
| 182 |
|
| 183 |
+
- The live retrieval path is FAISS dense search; BM25 and reciprocal rank fusion are documented as roadmap work until implemented in code.
|
| 184 |
+
- Fine-Tuned mode depends on an available hosted endpoint or compatible local inference environment.
|
| 185 |
+
- Streamlit session state is not durable across browser restarts; the separate FastAPI memory backend demonstrates the production direction.
|
| 186 |
+
- The system is educational and portfolio-grade; it is not legal, financial, investment, or compliance advice.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 187 |
|
| 188 |
+
## Roadmap
|
| 189 |
|
| 190 |
+
See [`ROADMAP.md`](ROADMAP.md) for the planned reliability, agentic architecture, governance, and production-hardening phases.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 191 |
|
| 192 |
+
## License and Contact
|
| 193 |
|
| 194 |
+
This repository is intended as an AI engineering portfolio and educational system. For questions, reach out through [GitHub](https://github.com/rakeshmadasaniai), [Hugging Face](https://huggingface.co/RakeshMadasani), or [LinkedIn](https://www.linkedin.com/in/rakesh-madasani-b217b71b0/).
|
ROADMAP.md
ADDED
|
@@ -0,0 +1,49 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Roadmap
|
| 2 |
+
|
| 3 |
+
This roadmap keeps the project moving toward a stronger production-grade Banking & Finance AI Agent without overclaiming what is already implemented.
|
| 4 |
+
|
| 5 |
+
## Phase 1 - Reliability
|
| 6 |
+
|
| 7 |
+
- Fill any missing evaluation tables with reproducible result snapshots.
|
| 8 |
+
- Add regression tests for high-risk banking, compliance, investment, and multilingual paths.
|
| 9 |
+
- Improve source-grounding checks and weak-retrieval fallbacks.
|
| 10 |
+
- Track p50, p95, and p99 latency instead of only averages.
|
| 11 |
+
- Add deterministic smoke tests for the Streamlit runtime startup path.
|
| 12 |
+
|
| 13 |
+
## Phase 2 - Agentic Architecture
|
| 14 |
+
|
| 15 |
+
- Formalize a Planner -> Executor -> Verifier loop.
|
| 16 |
+
- Standardize tool-use traces across all agentic modes.
|
| 17 |
+
- Add retry policy, max-step budgets, and clear stop conditions.
|
| 18 |
+
- Separate "answer directly" from "act with tools" using an explicit decision layer.
|
| 19 |
+
- Add durable task-state persistence for long-running autonomous workflows.
|
| 20 |
+
|
| 21 |
+
## Phase 3 - Governance
|
| 22 |
+
|
| 23 |
+
- Add compliance guardrails for AML, KYC, sanctions, investment-risk, and crisis scenarios.
|
| 24 |
+
- Add structured audit logs for tool calls, verification decisions, and fallback paths.
|
| 25 |
+
- Add escalation policy for unsupported legal, compliance, or investment-advice requests.
|
| 26 |
+
- Add risk scoring for answers that combine regulated finance and user-specific facts.
|
| 27 |
+
- Create reviewer-friendly model cards and dataset cards with limitations clearly stated.
|
| 28 |
+
|
| 29 |
+
## Phase 4 - Production Hardening
|
| 30 |
+
|
| 31 |
+
- Add Docker deployment for the Streamlit app and FastAPI memory backend.
|
| 32 |
+
- Add CI/CD for tests, linting, and evaluation smoke checks.
|
| 33 |
+
- Add API documentation for the memory backend.
|
| 34 |
+
- Add monitoring for latency, tool failure rates, retrieval misses, and user-facing errors.
|
| 35 |
+
- Add versioned releases and changelogs.
|
| 36 |
+
|
| 37 |
+
## Phase 5 - Retrieval Upgrade
|
| 38 |
+
|
| 39 |
+
- Add BM25 sparse retrieval.
|
| 40 |
+
- Add reciprocal rank fusion between FAISS dense retrieval and BM25 sparse retrieval.
|
| 41 |
+
- Add retrieval ablation tests to quantify dense-only versus hybrid retrieval quality.
|
| 42 |
+
- Add query rewriting for difficult regulatory and multilingual prompts.
|
| 43 |
+
|
| 44 |
+
## Phase 6 - Portfolio Distribution
|
| 45 |
+
|
| 46 |
+
- Publish a concise technical write-up explaining the system design and evaluation.
|
| 47 |
+
- Mirror the live Hugging Face Space state into GitHub branches or releases.
|
| 48 |
+
- Add demo clips or GIFs that show source cards, agent trace, uploads, and voice output.
|
| 49 |
+
- Keep README claims tied to committed code, result files, and reproducible scripts.
|
SYSTEM_DESIGN.md
ADDED
|
@@ -0,0 +1,118 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# System Design
|
| 2 |
+
|
| 3 |
+
## Problem
|
| 4 |
+
|
| 5 |
+
Banking and finance users need answers that are clear, grounded, and measurable. A normal chatbot can produce fluent text, but regulated domains require source visibility, confidence signals, evaluation, and careful handling of user-specific finance or compliance scenarios.
|
| 6 |
+
|
| 7 |
+
## Goals
|
| 8 |
+
|
| 9 |
+
- Provide a live AI workflow interface for banking and compliance questions.
|
| 10 |
+
- Ground answers in curated banking material and uploaded documents.
|
| 11 |
+
- Support multiple model paths and routing decisions.
|
| 12 |
+
- Track latency, confidence, source usage, and evaluation outputs.
|
| 13 |
+
- Demonstrate the full system chain: data, retrieval, fine-tuning, memory, evaluation, and deployment.
|
| 14 |
+
|
| 15 |
+
## Non-Goals
|
| 16 |
+
|
| 17 |
+
- Do not claim AGI.
|
| 18 |
+
- Do not claim formal legal, investment, financial, or compliance advice.
|
| 19 |
+
- Do not claim BM25/RRF as implemented until the code path exists.
|
| 20 |
+
- Do not claim fully autonomous production operation without durable planning, monitoring, governance, and external action controls.
|
| 21 |
+
|
| 22 |
+
## Architecture
|
| 23 |
+
|
| 24 |
+
```mermaid
|
| 25 |
+
flowchart TD
|
| 26 |
+
A["User input"] --> B["Input router"]
|
| 27 |
+
B --> C["Text / upload / voice handling"]
|
| 28 |
+
C --> D["Chunking + embeddings"]
|
| 29 |
+
D --> E["FAISS dense retrieval"]
|
| 30 |
+
E --> F["Grounded context"]
|
| 31 |
+
F --> G["Model orchestration"]
|
| 32 |
+
G --> H["OpenAI mode"]
|
| 33 |
+
G --> I["Fine-Tuned mode"]
|
| 34 |
+
G --> J["Auto mode"]
|
| 35 |
+
G --> K["Agentic / Autonomous modes"]
|
| 36 |
+
H --> L["Confidence + source rendering"]
|
| 37 |
+
I --> L
|
| 38 |
+
J --> L
|
| 39 |
+
K --> L
|
| 40 |
+
L --> M["Response UI"]
|
| 41 |
+
M --> N["Session memory / audit context"]
|
| 42 |
+
```
|
| 43 |
+
|
| 44 |
+
## Retrieval Design
|
| 45 |
+
|
| 46 |
+
Current implementation:
|
| 47 |
+
|
| 48 |
+
- Documents are chunked and embedded.
|
| 49 |
+
- FAISS provides dense vector search.
|
| 50 |
+
- Retrieved context is passed into answer generation and source cards.
|
| 51 |
+
|
| 52 |
+
Planned retrieval hardening:
|
| 53 |
+
|
| 54 |
+
- Add BM25 sparse retrieval for exact regulatory terms.
|
| 55 |
+
- Fuse dense and sparse results with reciprocal rank fusion.
|
| 56 |
+
- Add retrieval evaluation comparing dense-only and hybrid retrieval.
|
| 57 |
+
|
| 58 |
+
## Model Orchestration
|
| 59 |
+
|
| 60 |
+
The runtime supports several paths:
|
| 61 |
+
|
| 62 |
+
- OpenAI mode for stable general-purpose answers.
|
| 63 |
+
- Fine-Tuned mode for the banking-domain adapter path.
|
| 64 |
+
- Auto mode for scoring candidate answers and selecting a winner.
|
| 65 |
+
- Agentic/autonomous modes for tool-style workflows and execution traces where configured.
|
| 66 |
+
|
| 67 |
+
Candidate scoring uses groundedness, completeness, and latency signals. This makes routing inspectable instead of hidden.
|
| 68 |
+
|
| 69 |
+
## Memory Design
|
| 70 |
+
|
| 71 |
+
The live Streamlit runtime uses session state for chat/session continuity during a browser session. The separate FastAPI memory backend demonstrates the production direction:
|
| 72 |
+
|
| 73 |
+
- session IDs,
|
| 74 |
+
- retained recent turns,
|
| 75 |
+
- summarization/truncation,
|
| 76 |
+
- comparison endpoints,
|
| 77 |
+
- health checks.
|
| 78 |
+
|
| 79 |
+
Production memory should move to durable storage such as Redis or Postgres.
|
| 80 |
+
|
| 81 |
+
## Evaluation Design
|
| 82 |
+
|
| 83 |
+
Evaluation is repository-native:
|
| 84 |
+
|
| 85 |
+
- domain prompt packs,
|
| 86 |
+
- multilingual prompt packs,
|
| 87 |
+
- runner scripts,
|
| 88 |
+
- summarizer scripts,
|
| 89 |
+
- committed CSV/JSON outputs,
|
| 90 |
+
- generated portfolio report,
|
| 91 |
+
- autonomy audit.
|
| 92 |
+
|
| 93 |
+
This structure lets reviewers inspect prompts and outputs rather than relying on hand-picked examples.
|
| 94 |
+
|
| 95 |
+
## Reliability Considerations
|
| 96 |
+
|
| 97 |
+
- Weak retrieval should be surfaced clearly.
|
| 98 |
+
- Mode failures should degrade gracefully.
|
| 99 |
+
- Fine-Tuned endpoint unavailability should not crash the app.
|
| 100 |
+
- Upload parsing should handle unsupported or malformed files safely.
|
| 101 |
+
- Agentic workflows need max-step limits, retry budgets, and stop conditions.
|
| 102 |
+
- Regulated finance answers should include disclaimers and avoid unsupported personalized advice.
|
| 103 |
+
|
| 104 |
+
## Tradeoffs
|
| 105 |
+
|
| 106 |
+
- Streamlit is excellent for fast product iteration, but not a full production backend by itself.
|
| 107 |
+
- FAISS dense retrieval is simple and fast, but sparse retrieval is needed for exact regulatory threshold matching.
|
| 108 |
+
- QLoRA adapters are efficient for domain adaptation, but inference still requires a compatible base model environment.
|
| 109 |
+
- Agentic behavior improves complex workflows, but it adds latency and requires strict controls.
|
| 110 |
+
|
| 111 |
+
## Future Improvements
|
| 112 |
+
|
| 113 |
+
- Implement BM25 + reciprocal rank fusion.
|
| 114 |
+
- Add durable memory and audit logs.
|
| 115 |
+
- Add Docker and CI/CD.
|
| 116 |
+
- Add p95/p99 latency monitoring.
|
| 117 |
+
- Add benchmarked document-upload and voice tests.
|
| 118 |
+
- Add governance guardrails for high-risk compliance and investment scenarios.
|
requirements.txt
CHANGED
|
@@ -1,12 +1,16 @@
|
|
|
|
|
| 1 |
langchain==0.1.20
|
| 2 |
langchain-community==0.0.38
|
| 3 |
langchain-openai==0.1.6
|
| 4 |
-
openai=
|
|
|
|
| 5 |
httpx<0.28
|
| 6 |
faiss-cpu
|
| 7 |
sentence-transformers==2.7.0
|
| 8 |
pypdf==4.2.0
|
|
|
|
| 9 |
python-dotenv==1.0.1
|
| 10 |
huggingface-hub==0.23.0
|
| 11 |
python-docx==1.1.2
|
| 12 |
streamlit-mic-recorder==0.0.8
|
|
|
|
|
|
| 1 |
+
streamlit==1.56.0
|
| 2 |
langchain==0.1.20
|
| 3 |
langchain-community==0.0.38
|
| 4 |
langchain-openai==0.1.6
|
| 5 |
+
openai>=1.40.0
|
| 6 |
+
pydantic>=2.0.0
|
| 7 |
httpx<0.28
|
| 8 |
faiss-cpu
|
| 9 |
sentence-transformers==2.7.0
|
| 10 |
pypdf==4.2.0
|
| 11 |
+
PyPDF2>=3.0.0
|
| 12 |
python-dotenv==1.0.1
|
| 13 |
huggingface-hub==0.23.0
|
| 14 |
python-docx==1.1.2
|
| 15 |
streamlit-mic-recorder==0.0.8
|
| 16 |
+
yfinance
|