rakeshmadasaniai commited on
Commit
7d48dea
·
1 Parent(s): ca99372

Polish AI agent portfolio documentation

Browse files
.env.example ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ OPENAI_API_KEY=
2
+ HF_TOKEN=
3
+ MODEL_MODE=OpenAI
4
+ TOP_K=4
5
+ TEMPERATURE=0.2
6
+
7
+ # Optional runtime configuration
8
+ OPENAI_MODEL=gpt-4o-mini
9
+ OPENAI_STT_MODEL=gpt-4o-mini-transcribe
10
+ OPENAI_TTS_MODEL=gpt-4o-mini-tts
11
+ OPENAI_TTS_VOICE=alloy
12
+ FINETUNED_MODEL_ID=RakeshMadasani/banking-finance-mistral-qlora
13
+ FINETUNED_ENDPOINT_URL=
.gitignore CHANGED
@@ -1,7 +1,72 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  .venv/
 
 
2
  **/.venv/
3
- **/__pycache__/
4
- *.pyc
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
5
  01-rag-system/screenshots/*.png
6
  01-rag-system/screenshots/*.jpg
7
  01-rag-system/screenshots/*.jpeg
 
1
+ # Python
2
+ __pycache__/
3
+ **/__pycache__/
4
+ *.py[cod]
5
+ *.pyo
6
+ *.pyd
7
+ .Python
8
+ *.egg-info/
9
+ .pytest_cache/
10
+ .mypy_cache/
11
+ .ruff_cache/
12
+ .coverage
13
+ htmlcov/
14
+
15
+ # Virtual environments
16
  .venv/
17
+ venv/
18
+ env/
19
  **/.venv/
20
+ **/venv/
21
+
22
+ # Environment and secrets
23
+ .env
24
+ .env.*
25
+ !.env.example
26
+ **/.env
27
+ **/.env.*
28
+ !**/.env.example
29
+ .streamlit/secrets.toml
30
+ **/.streamlit/secrets.toml
31
+
32
+ # Streamlit and local app state
33
+ .streamlit/
34
+ **/.streamlit/
35
+ *.log
36
+ logs/
37
+
38
+ # Local model artifacts and generated indexes
39
+ models/
40
+ checkpoints/
41
+ outputs/
42
+ wandb/
43
+ runs/
44
+ mlruns/
45
+ *.safetensors
46
+ *.bin
47
+ *.pt
48
+ *.pth
49
+ *.onnx
50
+ *.gguf
51
+ **/faiss_index/
52
+ **/vectorstore/
53
+ **/*.faiss
54
+ **/*.pkl
55
+
56
+ # Data artifacts that should not be committed accidentally
57
+ *.parquet
58
+ *.arrow
59
+ *.sqlite
60
+ *.db
61
+ *.duckdb
62
+
63
+ # OS and editor files
64
+ .DS_Store
65
+ Thumbs.db
66
+ .vscode/
67
+ .idea/
68
+
69
+ # Large local screenshots; curated tracked screenshots can be force-added if needed
70
  01-rag-system/screenshots/*.png
71
  01-rag-system/screenshots/*.jpg
72
  01-rag-system/screenshots/*.jpeg
01-rag-system/README.md CHANGED
@@ -1,194 +1,97 @@
1
- # Autonomous Banking & Finance AI Agent
2
 
3
- Banking & Finance Copilot is a grounded AI product for USA and India banking, compliance, and financial intelligence. It combines retrieval over curated banking material with multiple model paths, source-backed answer cards, uploads, multilingual support, and evaluation workflows that are committed in the repo alongside the app itself.
4
 
5
- ## Live Product
6
 
7
- [banking-finance-rag on Hugging Face](https://huggingface.co/spaces/RakeshMadasani/banking-finance-rag)
8
 
9
- ## What This Project Is
10
-
11
- This is the product layer of the broader portfolio. The goal was not just to make a banking chatbot answer questions. The goal was to make it behave like a product someone could open, test, trust, and discuss seriously:
12
-
13
- - grounded answers instead of free-floating generation
14
- - visible source support
15
- - multiple model modes with clear routing behavior
16
- - uploads for real user documents
17
- - multilingual interaction
18
- - reproducible evaluation, not just screenshots
19
-
20
- ## Where The Code Is
21
-
22
- The live Hugging Face Space is only the deployment target. The actual product implementation is committed here in this project:
23
-
24
- - [`core`](./core)
25
- retrieval orchestration, runtime flow, prompts, and shared utilities
26
- - [`features`](./features)
27
- UI rendering, uploads, read-aloud, answer formatting, and user interaction
28
- - [`models`](./models)
29
- OpenAI mode, Fine-Tuned mode, and Auto routing logic
30
-
31
- That is important because I wanted the repo to stand on its own as a real product codebase, not just point outward to a demo URL.
32
-
33
- ## What The User Can Do
34
-
35
- - ask banking, AML, KYC, FDIC, Basel III, RBI, and compliance questions
36
- - switch between `OpenAI`, `Fine-Tuned`, and `Auto` modes
37
- - upload PDF, DOCX, TXT, and image files
38
- - inspect retrieved sources under each answer
39
- - use read-aloud on the final response
40
- - test multilingual questions
41
-
42
- ## Why The Product Is Structured This Way
43
-
44
- Trust was the main design constraint. For finance and compliance questions, a polished answer alone is not enough. The product needs to show where the answer came from and make its behavior explainable.
45
-
46
- That is why the app is built around:
47
-
48
- - shared retrieval before generation
49
- - source cards and chunk previews
50
- - model-mode transparency
51
- - latency and confidence visibility
52
- - evaluation packs committed in the repository
53
-
54
- ## Model Modes
55
-
56
- ### OpenAI
57
-
58
- This is the strongest general-purpose answer path and the most stable baseline for live testing.
59
-
60
- ### Fine-Tuned
61
-
62
- This uses the banking-domain Mistral adapter. It is valuable when the hosted path is configured and when lower-latency or domain-style responses are desirable.
63
-
64
- ### Auto
65
-
66
- Auto retrieves once, evaluates candidate answer paths, and selects the winner. That makes the routing logic easier to reason about than a hidden black-box switch.
67
-
68
- ## Product Architecture
69
 
70
  ```mermaid
71
- flowchart LR
72
- A["Built-in banking knowledge + uploaded docs"] --> B["Chunking and preprocessing"]
73
- B --> C["Embeddings"]
74
- C --> D["FAISS retrieval"]
75
- D --> E["Shared grounded context"]
76
- E --> F["OpenAI mode"]
77
- E --> G["Fine-Tuned mode"]
78
- E --> H["Auto routing"]
79
- F --> I["Answer card"]
80
- G --> I
81
- H --> I
82
- I --> J["Sources, confidence, latency, read aloud"]
 
 
 
83
  ```
84
 
85
- ## How It Works
86
-
87
- 1. The app loads curated banking knowledge files and any uploaded user documents.
88
- 2. Documents are chunked and embedded.
89
- 3. FAISS retrieves the most relevant context for the question.
90
- 4. The selected model mode answers from that shared context.
91
- 5. The UI renders the answer together with:
92
- - mode
93
- - latency
94
- - chunk count
95
- - source cards
96
- - confidence label
97
-
98
- ## Product Walkthrough
99
-
100
- ### A clean first impression for the product
101
-
102
- This is the opening experience of the Banking & Finance Copilot: the stable sidebar, the mode selector, the welcome guidance, and multilingual starter questions that make the product feel usable from the first click.
103
-
104
- ![Banking Copilot home experience](screenshots/banking-copilot-home-experience.png)
105
-
106
- ### A grounded English answer that feels concise and useful
107
-
108
- This example shows the assistant answering a KYC question in English with a direct explanation, short supporting bullets, visible latency, and a retrieved source card underneath the answer.
109
-
110
- ![English KYC answer walkthrough](screenshots/english-kyc-answer-walkthrough.png)
111
-
112
- ### The same product experience working in Telugu
113
 
114
- This screenshot matters because it shows the product doing more than translation. The answer stays structured, readable, and grounded while responding to the question naturally in Telugu.
 
 
 
 
 
 
 
 
 
 
 
 
115
 
116
- ![Telugu KYC answer walkthrough](screenshots/telugu-kyc-answer-walkthrough.png)
117
 
118
- ### Multilingual grounding working in Chinese as well
119
-
120
- This example shows the same KYC flow in Chinese, which helps demonstrate that the product experience is consistent across languages rather than being strong only in English.
121
-
122
- ![Chinese KYC answer walkthrough](screenshots/chinese-kyc-answer-walkthrough.png)
123
-
124
- ## Evaluation
125
-
126
- The [`evaluation`](./evaluation) folder includes two larger committed evaluation packs:
127
-
128
- - `evaluation_queries.md`
129
- 120 domain-specific prompts across OpenAI, Fine-Tuned, and Auto
130
- - `evaluation_multilingual.md`
131
- 120 multilingual prompts across the same three modes
132
-
133
- Supporting scripts:
134
-
135
- - `run_eval_sets.py`
136
- - `summarize_eval_sets.py`
137
-
138
- Latest committed result snapshots live in [`evaluation/results`](./evaluation/results).
139
- Autonomy audit for current release is tracked in [`../AUTONOMY_EVALUATION.md`](../AUTONOMY_EVALUATION.md).
140
-
141
- ### Latest committed summaries
142
-
143
- | Evaluation set | Total prompts | Available evaluated rows | Average latency | Median latency |
144
- |---|---:|---:|---:|---:|
145
- | Domain set | 120 | 80 | 2037.0 ms | 2036.0 ms |
146
- | Multilingual set | 120 | 80 | 2031.8 ms | 2031.5 ms |
147
-
148
- ### Reading the snapshot correctly
149
-
150
- Those numbers are the committed run snapshot, not a made-up "best case" table:
151
 
152
- - each pack contains 120 prompts
153
- - 80 rows were available in the committed export
154
- - the missing rows reflect backend availability in that local run, not missing evaluation logic
155
 
156
- I prefer that level of honesty because anyone reviewing the repo can inspect the CSVs and JSON summaries directly and see what was measured versus what was unavailable in that environment.
 
 
 
 
 
157
 
158
- ## Run Locally
159
 
160
- ### Prerequisites
161
 
162
- - Python 3.10+
163
- - OpenAI API key
 
 
164
 
165
- ### Install
166
 
167
- ```bash
168
- pip install -r requirements.txt
169
- ```
 
 
170
 
171
- ### Environment
172
 
173
- ```bash
174
- OPENAI_API_KEY=your_api_key_here
175
- OPENAI_MODEL=gpt-4o-mini
176
- OPENAI_STT_MODEL=gpt-4o-mini-transcribe
177
- OPENAI_TTS_MODEL=gpt-4o-mini-tts
178
- OPENAI_TTS_VOICE=alloy
179
- FINETUNED_MODEL_ID=RakeshMadasani/banking-finance-mistral-qlora
180
- FINETUNED_ENDPOINT_URL=
181
- HF_TOKEN=your_hugging_face_token
182
- ```
183
 
184
- ### Start the app
 
 
 
 
 
185
 
186
- ```bash
187
- streamlit run app.py
188
- ```
189
 
190
- ## Notes
191
 
192
- - Fine-Tuned mode becomes fully live when a hosted endpoint is configured. The adapter, routing logic, and evaluation path are in the repo today; the hosted endpoint is the last operational piece for a fully public demo of that mode.
193
- - Voice input and some upload flows are still environment-sensitive because they depend on browser/runtime behavior.
194
- - The app is built for groundedness and explainability first, not raw throughput.
 
 
1
+ # AI Agent Runtime
2
 
3
+ This folder contains the live deployed Banking & Finance AI Agent runtime. It is still named `01-rag-system` for Hugging Face deployment compatibility, but its role in the system is the agent runtime: Streamlit UI, retrieval, model routing, agentic workflows, uploads, voice controls, source cards, and evaluation.
4
 
5
+ ## Purpose
6
 
7
+ Provide a product-quality AI workflow interface for banking, compliance, and financial knowledge tasks. The runtime prioritizes grounded answers, clear confidence signals, source visibility, multilingual support, and stable user interaction.
8
 
9
+ ## Architecture
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
10
 
11
  ```mermaid
12
+ flowchart TD
13
+ A["User input"] --> B["Streamlit product runtime"]
14
+ B --> C["Upload / voice / text handling"]
15
+ C --> D["Chunking and embeddings"]
16
+ D --> E["FAISS dense retrieval"]
17
+ E --> F["Grounded context"]
18
+ F --> G["OpenAI mode"]
19
+ F --> H["Fine-Tuned mode"]
20
+ F --> I["Auto mode"]
21
+ F --> J["Agentic / Autonomous modes"]
22
+ G --> K["Answer renderer"]
23
+ H --> K
24
+ I --> K
25
+ J --> K
26
+ K --> L["Sources, confidence, latency, actions, read aloud"]
27
  ```
28
 
29
+ ## Key Files
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
30
 
31
+ | File or folder | Role |
32
+ |---|---|
33
+ | `app.py` | Streamlit entry point used by the live Hugging Face Space. |
34
+ | `core/product_runtime.py` | Main product orchestration, mode selection, retrieval calls, session state, and response handling. |
35
+ | `core/agentic_runtime.py` | Agentic/autonomous workflow implementation and tool-style reasoning layer. |
36
+ | `core/retriever.py` | Shared context retrieval over runtime indexes. |
37
+ | `core/vector_store.py` | FAISS vector store construction. |
38
+ | `features/product_ui.py` | Premium UI cards, sidebar, metrics, and answer rendering. |
39
+ | `features/voice_input.py` / `features/voice_output.py` | Speech-to-text and text-to-speech integration paths. |
40
+ | `models/openai_mode.py` | OpenAI answer path. |
41
+ | `models/finetuned_mode.py` | Fine-tuned model endpoint path. |
42
+ | `models/auto_router.py` | Candidate scoring and automatic model selection. |
43
+ | `evaluation/` | Domain and multilingual evaluation packs, runners, summaries, and reports. |
44
 
45
+ ## How To Run
46
 
47
+ ```bash
48
+ cd 01-rag-system
49
+ pip install -r requirements.txt
50
+ streamlit run app.py
51
+ ```
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
52
 
53
+ Required environment:
 
 
54
 
55
+ ```env
56
+ OPENAI_API_KEY=
57
+ HF_TOKEN=
58
+ OPENAI_MODEL=gpt-4o-mini
59
+ FINETUNED_MODEL_ID=RakeshMadasani/banking-finance-mistral-qlora
60
+ ```
61
 
62
+ ## Inputs And Outputs
63
 
64
+ Inputs:
65
 
66
+ - User questions in English or supported multilingual prompts.
67
+ - PDF, DOCX, TXT, and image-oriented upload workflows.
68
+ - Voice input where browser/runtime support is available.
69
+ - Mode selection across OpenAI, Fine-Tuned, Auto, Agentic Workspace, and Autonomous Max.
70
 
71
+ Outputs:
72
 
73
+ - Source-grounded answer cards.
74
+ - Confidence label and latency.
75
+ - Retrieved chunk/source metadata.
76
+ - Copy/export actions and read-aloud audio.
77
+ - Agent trace and audit-style context where agentic modes are used.
78
 
79
+ ## Evaluation Notes
80
 
81
+ The runtime includes:
 
 
 
 
 
 
 
 
 
82
 
83
+ - 120 domain prompts in `evaluation/evaluation_queries.md`.
84
+ - 120 multilingual prompts in `evaluation/evaluation_multilingual.md`.
85
+ - Runner and summarizer scripts for repeatable evaluation.
86
+ - Committed snapshots under `evaluation/results`.
87
+ - A generated report under `evaluation/reports/latest_portfolio_report.md`.
88
+ - Decision-critical tests under `tests/`.
89
 
90
+ Latest committed snapshots show about 2.03s average latency for available rows in both domain and multilingual evaluation exports.
 
 
91
 
92
+ ## Limitations
93
 
94
+ - Current retrieval is FAISS dense vector retrieval. BM25 and reciprocal rank fusion are roadmap work unless implemented later.
95
+ - Fine-Tuned mode requires a configured endpoint or compatible local runtime.
96
+ - Streamlit session state is suitable for live demos but not durable production memory by itself.
97
+ - Outputs are educational and must not be treated as legal, investment, financial, or compliance advice.
02-qa-dataset/README.md CHANGED
@@ -1,69 +1,87 @@
1
- # Banking & Finance QA Dataset
2
 
3
- This project contains the banking and finance instruction dataset used for downstream fine-tuning.
 
 
 
 
4
 
5
  ## Dataset
6
- [banking-finance-qa-dataset](https://huggingface.co/datasets/RakeshMadasani/banking-finance-qa-dataset)
7
 
8
- ## Screenshot
9
- Published dataset page with train and validation splits:
10
 
11
  ![Dataset on Hugging Face](screenshots/dataset-hf-splits.png)
12
 
13
  ## Summary
14
- - 3,002 instruction-response pairs
15
- - 2,701 train samples
16
- - 301 validation samples
17
- - Alpaca-style format
18
- - English language
19
- - Banking and finance domain
 
 
 
 
 
 
 
 
 
 
 
 
 
 
20
 
21
  ## Coverage
22
- - AML / KYC
23
- - CDD / EDD
24
- - FDIC
25
- - Basel III
26
- - RBI
27
- - SAR / CTR
28
- - compliance topics
29
- - finance fundamentals
30
-
31
- ## What it demonstrates
32
- - domain data curation
33
- - instruction dataset creation
34
- - validation and split design
35
- - Hugging Face dataset publishing
36
-
37
- ## Data Methodology
38
-
39
- The dataset was built by turning curated banking and compliance material into instruction-style examples intended for downstream fine-tuning. The focus was not just on volume, but on having a usable spread of topics, clean formatting, and a validation pass before publishing.
40
-
41
- In practice, that meant:
42
-
43
- - selecting questions that map naturally to banking, compliance, and regulatory supervision
44
- - structuring examples in Alpaca-style `instruction`, `input`, and `output` fields
45
- - validating the finished dataset for structural consistency and duplicate issues before upload
46
-
47
- ## Example Schema
48
-
49
- ```json
50
- {
51
- "instruction": "What is the FDIC deposit insurance limit in the United States?",
52
- "input": "",
53
- "output": "The FDIC insures deposits up to $250,000 per depositor, per insured bank, per account ownership category."
54
- }
55
  ```
56
 
57
- ## Code Entry Points
58
- - `generate_dataset.py` - builds banking instruction-response pairs from curated source material
59
- - `validate_dataset.py` - runs duplicate and structural checks on the generated dataset
60
- - `upload_to_hf.py` - publishes the dataset and dataset card to Hugging Face
 
 
 
 
 
 
 
 
 
 
 
 
61
 
62
- ## Why it stands out
63
- This project turns raw banking and compliance material into a reusable ML asset rather than stopping at prompt experimentation. It shows the data layer behind the model and application work in the rest of the portfolio.
64
 
65
  ## Limitations
66
 
67
- - instruction-style datasets can still inherit phrasing bias from the source material used to create them
68
- - U.S. and India banking topics are useful together, but not always perfectly balanced in representation
69
- - the dataset is a strong domain adaptation asset, but it should not be treated as formal regulatory ground truth
 
1
+ # BankingQA Dataset
2
 
3
+ This folder contains the BankingQA-3K instruction dataset workflow used by the Banking & Finance AI Agent.
4
+
5
+ ## Purpose
6
+
7
+ Create a reusable domain dataset for banking and financial compliance instruction tuning. The dataset turns curated banking material into structured instruction-response pairs that support the downstream QLoRA adaptation workflow.
8
 
9
  ## Dataset
 
10
 
11
+ [banking-finance-qa-dataset](https://huggingface.co/datasets/RakeshMadasani/banking-finance-qa-dataset)
 
12
 
13
  ![Dataset on Hugging Face](screenshots/dataset-hf-splits.png)
14
 
15
  ## Summary
16
+
17
+ | Item | Value |
18
+ |---|---:|
19
+ | Total examples | 3,002 |
20
+ | Train samples | 2,701 |
21
+ | Validation samples | 301 |
22
+ | Format | Alpaca-style instruction data |
23
+ | Language | English |
24
+ | Domain | Banking, finance, AML, KYC, compliance |
25
+
26
+ ## Architecture
27
+
28
+ ```mermaid
29
+ flowchart LR
30
+ A["Curated banking material"] --> B["Question generation"]
31
+ B --> C["Instruction / input / output schema"]
32
+ C --> D["Validation and duplicate checks"]
33
+ D --> E["Train / validation split"]
34
+ E --> F["Hugging Face Dataset"]
35
+ ```
36
 
37
  ## Coverage
38
+
39
+ - AML, KYC, CDD, and EDD.
40
+ - FDIC deposit insurance.
41
+ - Basel III capital concepts.
42
+ - RBI and India banking compliance topics.
43
+ - SAR, CTR, transaction monitoring, and financial crime concepts.
44
+ - General banking and finance fundamentals.
45
+
46
+ ## Key Files
47
+
48
+ | File | Role |
49
+ |---|---|
50
+ | `generate_dataset.py` | Builds instruction-response examples from curated material. |
51
+ | `validate_dataset.py` | Checks structure, duplicates, and dataset quality signals. |
52
+ | `upload_to_hf.py` | Publishes dataset artifacts and dataset card to Hugging Face. |
53
+ | `screenshots/` | Published dataset page screenshots for portfolio review. |
54
+
55
+ ## How To Run
56
+
57
+ ```bash
58
+ cd 02-qa-dataset
59
+ python generate_dataset.py
60
+ python validate_dataset.py
61
+ python upload_to_hf.py
 
 
 
 
 
 
 
 
 
62
  ```
63
 
64
+ Set `HF_TOKEN` before upload if publishing to the Hub.
65
+
66
+ ## Inputs And Outputs
67
+
68
+ Inputs:
69
+
70
+ - Curated banking and compliance source material.
71
+ - Topic coverage plan across AML, KYC, FDIC, RBI, Basel III, and banking operations.
72
+
73
+ Outputs:
74
+
75
+ - Alpaca-style dataset with `instruction`, `input`, and `output` fields.
76
+ - Train and validation splits.
77
+ - Hugging Face dataset repository.
78
+
79
+ ## Evaluation Notes
80
 
81
+ The dataset is validated structurally before publishing. It is designed for fine-tuning and domain adaptation, not as a formal legal or regulatory authority. Downstream quality should be measured through model evaluation and grounded answer testing.
 
82
 
83
  ## Limitations
84
 
85
+ - The dataset is English-only in its current published form.
86
+ - Coverage is strongest for banking/compliance concepts represented in the curated material.
87
+ - Dataset answers should be treated as training material, not official regulatory advice.
03-qlora-finetuning/README.md CHANGED
@@ -1,79 +1,83 @@
1
- # Banking Finance QLoRA Fine-Tuned Model
2
 
3
- This project contains the fine-tuning workflow for adapting a Mistral-based LLM to banking and finance question answering.
 
 
 
 
4
 
5
  ## Model
 
6
  [banking-finance-mistral-qlora](https://huggingface.co/RakeshMadasani/banking-finance-mistral-qlora)
7
 
8
- ## Published Model Page
9
  ![Published QLoRA model page](screenshots/model-page-demo.png)
10
 
11
- ## Recommended Demo Questions
12
 
13
- If you want to capture a stronger inference screenshot for this project, use these questions:
 
 
 
 
 
 
 
 
14
 
15
- - `What is the FDIC deposit insurance limit in the United States?`
16
- - `What are the three stages of money laundering?`
17
- - `What is the difference between AML and KYC?`
18
-
19
- These prompts are short, easy to judge, and representative of the banking/compliance domain adaptation shown by the model.
20
 
21
- ## Base Model
22
- `mistralai/Mistral-7B-Instruct-v0.3`
 
 
 
 
 
 
 
 
 
 
23
 
24
- ## Fine-Tuning Summary
25
- - Method: QLoRA
26
- - Quantization: 4-bit NF4
27
- - LoRA rank: 16
28
- - LoRA alpha: 32
29
- - LoRA dropout: 0.05
30
- - Training samples: 2,701
31
- - Validation samples: 301
32
- - Global steps: 676
33
- - Final train loss: 1.13
34
 
35
- ## Training Environment
 
 
 
 
36
 
37
- - notebook-based QLoRA workflow
38
- - designed for practical notebook GPU usage rather than full distributed training
39
- - built around an adapter approach so a 7B model could be adapted without full fine-tuning
40
 
41
- ## Code Entry Points
42
- - `Banking_QLoRA_Mistral7B_updated.ipynb` - main notebook used for the QLoRA workflow
43
- - `inference_demo.py` - lightweight inference script for loading the published adapter and testing example prompts
44
 
45
- ## Why QLoRA
 
 
 
46
 
47
- - **4-bit NF4 quantization:** reduces memory usage enough to make 7B-scale fine-tuning practical in notebook GPU environments
48
- - **LoRA adapters:** updates a small trainable parameter set instead of full-model weights
49
- - **Cost-efficient experimentation:** a good fit for domain adaptation when full fine-tuning is too heavy
50
 
51
- ## What it demonstrates
52
- - parameter-efficient fine-tuning
53
- - PEFT/LoRA configuration
54
- - domain adaptation using custom data
55
- - Hugging Face model publishing
56
 
57
- ## Output Snapshot
58
 
59
- Example sample outputs from the fine-tuned model included banking-domain answers for prompts such as:
60
- - FDIC deposit insurance limits
61
- - the stages of money laundering
62
- - Basel-related banking questions
63
 
64
- These examples showed that the adapter was capable of producing domain-specific responses after fine-tuning, even though formal benchmark scoring is still a planned improvement.
65
 
66
- ## Before vs After Snapshot
 
 
67
 
68
- | Prompt type | Base model tendency | Fine-tuned model tendency |
69
- |---|---|---|
70
- | Banking definitions | usually reasonable but generic | more direct banking-specific phrasing |
71
- | Compliance language | broad but sometimes high-level | more targeted AML / KYC / SAR / CTR terminology |
72
- | India-focused regulation | can be vague or mix jurisdictions | better alignment with RBI-oriented wording from the custom dataset |
73
 
74
- ## Why it stands out
75
- This project shows that the portfolio goes beyond app development and dataset creation into actual model adaptation. Publishing the adapter with a model card, tokenizer files, and LoRA configuration makes the work visible and inspectable as a real model artifact.
76
 
77
- ## Limitation to state clearly
78
 
79
- This repository publishes a QLoRA adapter artifact, not a fully hosted standalone inference service by itself. To run it end to end, the adapter still needs to be loaded with the compatible base model in a suitable inference environment.
 
 
 
1
+ # Domain Model Adaptation
2
 
3
+ This folder contains the QLoRA fine-tuning workflow used to adapt Mistral-7B-Instruct-v0.3 to banking and financial compliance terminology.
4
+
5
+ ## Purpose
6
+
7
+ Show the model adaptation layer behind the Banking & Finance AI Agent. The goal is not only to call external APIs, but to demonstrate data preparation, parameter-efficient fine-tuning, adapter publishing, and inference testing.
8
 
9
  ## Model
10
+
11
  [banking-finance-mistral-qlora](https://huggingface.co/RakeshMadasani/banking-finance-mistral-qlora)
12
 
 
13
  ![Published QLoRA model page](screenshots/model-page-demo.png)
14
 
15
+ ## Architecture
16
 
17
+ ```mermaid
18
+ flowchart LR
19
+ A["BankingQA-3K dataset"] --> B["Prompt formatting"]
20
+ B --> C["Mistral-7B-Instruct-v0.3"]
21
+ C --> D["4-bit NF4 quantization"]
22
+ D --> E["LoRA adapter training"]
23
+ E --> F["Validation / training metrics"]
24
+ F --> G["Published Hugging Face adapter"]
25
+ ```
26
 
27
+ ## Fine-Tuning Summary
 
 
 
 
28
 
29
+ | Item | Value |
30
+ |---|---|
31
+ | Base model | `mistralai/Mistral-7B-Instruct-v0.3` |
32
+ | Method | QLoRA |
33
+ | Quantization | 4-bit NF4 |
34
+ | LoRA rank | 16 |
35
+ | LoRA alpha | 32 |
36
+ | LoRA dropout | 0.05 |
37
+ | Training samples | 2,701 |
38
+ | Validation samples | 301 |
39
+ | Global steps | 676 |
40
+ | Final train loss | 1.13 |
41
 
42
+ ## Key Files
 
 
 
 
 
 
 
 
 
43
 
44
+ | File | Role |
45
+ |---|---|
46
+ | `Banking_QLoRA_Mistral7B_updated.ipynb` | Main notebook for dataset formatting, QLoRA setup, training, and publishing. |
47
+ | `inference_demo.py` | Lightweight script for loading the adapter and testing domain prompts. |
48
+ | `screenshots/` | Published model page and training-progress screenshots. |
49
 
50
+ ## How To Run
 
 
51
 
52
+ The notebook is the main training artifact. For local inference, use:
 
 
53
 
54
+ ```bash
55
+ cd 03-qlora-finetuning
56
+ python inference_demo.py
57
+ ```
58
 
59
+ You need access to the base model, the published adapter, and a compatible local GPU/CPU environment. Full 7B inference can be heavy on consumer machines.
 
 
60
 
61
+ ## Inputs And Outputs
 
 
 
 
62
 
63
+ Inputs:
64
 
65
+ - BankingQA-3K instruction dataset.
66
+ - Mistral-7B-Instruct-v0.3 base model.
67
+ - PEFT/QLoRA configuration.
 
68
 
69
+ Outputs:
70
 
71
+ - Published LoRA adapter.
72
+ - Training metrics.
73
+ - Model card and inference demo path.
74
 
75
+ ## Evaluation Notes
 
 
 
 
76
 
77
+ The current folder documents training configuration and final train loss. A stronger future benchmark should compare the base model, adapter model, OpenAI path, and Auto routing on the same held-out banking evaluation pack.
 
78
 
79
+ ## Limitations
80
 
81
+ - This publishes an adapter artifact, not a fully hosted standalone inference service.
82
+ - Runtime quality depends on loading the compatible base model plus adapter correctly.
83
+ - The model should be evaluated on held-out prompts before making strong accuracy claims.
04-conversational-memory/README.md CHANGED
@@ -1,139 +1,51 @@
1
- # Conversational Memory Backend
2
 
3
- This project adds a FastAPI-based conversational layer on top of the banking RAG assistant so interactions become session-aware instead of stateless. It introduces session memory, controlled history retention, and summarization for longer conversations while keeping the underlying retrieval-backed banking Q&A workflow reusable. The backend supports both an OpenAI path for quick deployment and a local Hugging Face adapter path for a stronger end-to-end portfolio story.
4
 
5
- ## What this project does
6
 
7
- - wraps the existing banking RAG logic behind a FastAPI service
8
- - maintains per-session conversation history for multi-turn follow-up questions
9
- - truncates and summarizes long sessions to control prompt growth
10
- - exposes clean API endpoints for chat, session clearing, and health checks
11
- - supports switchable LLM backends via environment variable
12
- - supports side-by-side backend comparison on the same retrieved context
13
- - supports evaluation of memory-on vs memory-off coherence
14
 
15
  ## Architecture
16
 
17
  ```mermaid
18
  flowchart LR
19
- A["Client / frontend"] --> B["POST /chat"]
20
- A --> B2["POST /chat/compare"]
21
- B --> C["Session memory store"]
22
- B2 --> C
23
- C --> D["Truncation or summarization"]
24
- D --> E["Retriever + shared banking knowledge base"]
25
- E --> F["Shared retrieved context"]
26
- F --> G["Single backend path"]
27
- F --> H["OpenAI backend"]
28
- F --> I["Local HF adapter backend"]
29
- G --> J["Single response"]
30
- H --> K["OpenAI answer"]
31
- I --> L["HF answer"]
32
- K --> M["Side-by-side comparison output"]
33
- L --> M
34
-
35
- N["DELETE /session/{session_id}"] --> C
36
- O["GET /health"] --> B
37
  ```
38
 
39
- ## Project Structure
40
-
41
- ```text
42
- 04-conversational-memory/
43
- |-- app/
44
- | |-- __init__.py
45
- | |-- main.py
46
- | |-- memory.py
47
- | |-- models.py
48
- | |-- rag_chain.py
49
- | `-- summarizer.py
50
- |-- evaluation/
51
- | |-- coherence_eval.py
52
- | `-- questions.csv
53
- |-- tests/
54
- | `-- test_memory.py
55
- |-- .env.example
56
- |-- README.md
57
- `-- requirements.txt
58
- ```
59
-
60
- ## API Endpoints
61
-
62
- ### `POST /chat`
63
- Accepts a user message plus an optional `session_id` and returns:
64
-
65
- - `session_id`
66
- - `response`
67
- - `turn_count`
68
- - `sources`
69
- - `confidence`
70
- - `history_used`
71
- - `summary_used`
72
-
73
- ### `DELETE /session/{session_id}`
74
- Clears the stored memory for a given session.
75
-
76
- ### `GET /health`
77
- Returns a lightweight health status and current session count.
78
-
79
- ### `POST /chat/compare`
80
- Runs both backends on the same question and retrieved context, then returns both answers side by side for direct comparison.
81
 
82
- ## How to run
 
 
 
 
 
 
 
 
83
 
84
- ### Prerequisites
85
-
86
- - Python 3.10+
87
- - banking knowledge files available in `01-rag-system/`
88
- - one backend configured:
89
- - `openai` with `OPENAI_API_KEY`
90
- - `local_hf` with model + adapter access and enough local GPU/compute
91
-
92
- ### Install
93
 
94
  ```bash
 
95
  pip install -r requirements.txt
96
- ```
97
-
98
- ### Environment
99
-
100
- ```bash
101
- LLM_BACKEND=openai
102
- OPENAI_API_KEY=your_openai_api_key_here
103
- OPENAI_MODEL=gpt-4o-mini
104
- LOCAL_HF_MODEL_ID=mistralai/Mistral-7B-Instruct-v0.3
105
- LOCAL_HF_ADAPTER_ID=RakeshMadasani/banking-finance-mistral-qlora
106
- LOCAL_HF_DEVICE=auto
107
- ```
108
-
109
- ### Backend options
110
-
111
- #### Option A - OpenAI backend
112
-
113
- ```bash
114
- LLM_BACKEND=openai
115
- OPENAI_API_KEY=your_openai_api_key_here
116
- OPENAI_MODEL=gpt-4o-mini
117
- ```
118
-
119
- #### Option B - Local HF backend
120
-
121
- ```bash
122
- LLM_BACKEND=local_hf
123
- LOCAL_HF_MODEL_ID=mistralai/Mistral-7B-Instruct-v0.3
124
- LOCAL_HF_ADAPTER_ID=RakeshMadasani/banking-finance-mistral-qlora
125
- LOCAL_HF_DEVICE=auto
126
- ```
127
-
128
- This path is stronger for the portfolio because it links Project 3 directly into Project 4, but it also requires a compatible local environment and enough compute to load the base model plus adapter.
129
-
130
- ### Run the API
131
-
132
- ```bash
133
  uvicorn app.main:app --reload --port 8000
134
  ```
135
 
136
- ### Example request
137
 
138
  ```bash
139
  curl -X POST "http://127.0.0.1:8000/chat" ^
@@ -141,88 +53,29 @@ curl -X POST "http://127.0.0.1:8000/chat" ^
141
  -d "{\"message\":\"What is KYC?\",\"session_id\":\"demo-session\",\"use_memory\":true}"
142
  ```
143
 
144
- ### Example comparison request
145
-
146
- ```bash
147
- curl -X POST "http://127.0.0.1:8000/chat/compare" ^
148
- -H "Content-Type: application/json" ^
149
- -d "{\"message\":\"What is the difference between AML and KYC?\",\"session_id\":\"compare-session\",\"use_memory\":true}"
150
- ```
151
-
152
- ## Evaluation
153
-
154
- The `evaluation/coherence_eval.py` script compares memory-off and memory-on responses on multi-turn conversations such as:
155
-
156
- - KYC followed by follow-up questions about its components and its relation to AML
157
- - Basel III followed by capital requirement and Basel II comparison prompts
158
- - SAR followed by reporting-threshold clarification
159
-
160
- Run:
161
-
162
- ```bash
163
- python evaluation/coherence_eval.py
164
- ```
165
-
166
- This prints:
167
-
168
- - average coherence score with memory off
169
- - average coherence score with memory on
170
- - relative improvement
171
-
172
- Use the resulting percentage as evidence for a claim like `17% improvement` only when your real run produces that number.
173
-
174
- ## Execution Proof
175
-
176
- This project has already been exercised locally at a lightweight level:
177
-
178
- - `/health` returned a successful status response
179
- - `/chat` returned a real answer
180
- - a follow-up `/chat` call on the same session showed `history_used = true`
181
- - `/chat/compare` is implemented and reachable, with the local HF runtime path dependent on the local model environment
182
-
183
- For the current execution record, see [`evaluation/results.md`](./evaluation/results.md).
184
-
185
- ## Design Decisions
186
-
187
- ### Why truncation + summarization instead of only windowing?
188
-
189
- - simple windowing drops older context entirely
190
- - summarization preserves earlier intent and constraints
191
- - retaining the most recent turns keeps the API responsive for active follow-ups
192
- - current session store is in-process for demo simplicity; a production upgrade path would be Redis or Postgres
193
-
194
- ### Why reuse the Project 1 knowledge base?
195
-
196
- - keeps the backend directly tied to your deployed banking RAG system
197
- - strengthens the portfolio story across app layer and backend layer
198
- - makes evaluation easier because the data and retrieval domain stay consistent
199
-
200
- ### Why support switchable backends?
201
-
202
- - OpenAI is the fastest path to ship and demo the architecture
203
- - a local HF backend creates a stronger data -> model -> backend portfolio chain
204
- - the same FastAPI layer can serve both paths without changing the API contract
205
 
206
- ### Why add `/chat/compare`?
207
 
208
- - it makes backend tradeoffs visible instead of hidden behind a config switch
209
- - it lets you compare a general-purpose model and a domain-adapted model on the same retrieved evidence
210
- - it creates a stronger portfolio talking point than a simple backend toggle
 
211
 
212
- ## What this adds to the portfolio
213
 
214
- This project upgrades the overall portfolio from a collection of artifacts into a more production-oriented system story:
 
 
 
 
215
 
216
- - Project 1 provides the deployed RAG interface
217
- - Project 2 provides the banking instruction dataset
218
- - Project 3 provides the fine-tuned model artifact
219
- - Project 4 adds the backend memory and orchestration layer for multi-turn interaction
220
 
221
- ## Execution Notes
222
 
223
- See [`evaluation/results.md`](./evaluation/results.md) for a lightweight execution record covering:
224
 
225
- - successful local `/health` response
226
- - successful local `/chat` response
227
- - successful memory follow-up showing `history_used = true`
228
- - the current state of `/chat/compare` validation
 
1
+ # Memory And Orchestration Backend
2
 
3
+ This folder contains a FastAPI conversational memory backend for the Banking & Finance AI Agent. It adds session handling, controlled history retention, summarization, and backend comparison around the same banking knowledge workflow.
4
 
5
+ ## Purpose
6
 
7
+ Move the system beyond stateless single-turn answers by introducing reusable API endpoints and memory-aware orchestration. This layer is separate from the live Streamlit Space so it can evolve toward production API deployment without destabilizing the public demo.
 
 
 
 
 
 
8
 
9
  ## Architecture
10
 
11
  ```mermaid
12
  flowchart LR
13
+ A["Client or frontend"] --> B["POST /chat"]
14
+ A --> C["POST /chat/compare"]
15
+ B --> D["Session memory store"]
16
+ C --> D
17
+ D --> E["History truncation / summarization"]
18
+ E --> F["Shared banking retrieval"]
19
+ F --> G["OpenAI backend"]
20
+ F --> H["Local HF adapter backend"]
21
+ G --> I["Response"]
22
+ H --> J["Comparison response"]
23
+ K["DELETE /session/{id}"] --> D
24
+ L["GET /health"] --> B
 
 
 
 
 
 
25
  ```
26
 
27
+ ## Key Files
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
28
 
29
+ | File | Role |
30
+ |---|---|
31
+ | `app/main.py` | FastAPI application and route definitions. |
32
+ | `app/memory.py` | Session memory store and state handling. |
33
+ | `app/rag_chain.py` | Retrieval and backend generation logic. |
34
+ | `app/summarizer.py` | Conversation summarization support. |
35
+ | `app/models.py` | Request and response models. |
36
+ | `evaluation/coherence_eval.py` | Memory-on versus memory-off coherence evaluation. |
37
+ | `tests/test_memory.py` | Unit tests for memory behavior. |
38
 
39
+ ## How To Run
 
 
 
 
 
 
 
 
40
 
41
  ```bash
42
+ cd 04-conversational-memory
43
  pip install -r requirements.txt
44
+ copy .env.example .env
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
45
  uvicorn app.main:app --reload --port 8000
46
  ```
47
 
48
+ Example:
49
 
50
  ```bash
51
  curl -X POST "http://127.0.0.1:8000/chat" ^
 
53
  -d "{\"message\":\"What is KYC?\",\"session_id\":\"demo-session\",\"use_memory\":true}"
54
  ```
55
 
56
+ ## Inputs And Outputs
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
57
 
58
+ Inputs:
59
 
60
+ - User message.
61
+ - Optional `session_id`.
62
+ - Optional memory usage flag.
63
+ - Backend configuration via environment variables.
64
 
65
+ Outputs:
66
 
67
+ - Session-aware answer.
68
+ - Sources and confidence.
69
+ - Turn count.
70
+ - Flags showing whether history or summary was used.
71
+ - Side-by-side backend comparison through `/chat/compare`.
72
 
73
+ ## Evaluation Notes
 
 
 
74
 
75
+ The coherence evaluation compares memory-off and memory-on behavior across follow-up conversations. Claims such as percentage improvement should only be made from actual generated evaluation output, not assumed values.
76
 
77
+ ## Limitations
78
 
79
+ - The current memory store is lightweight and in-process; production use should move to Redis, Postgres, or another durable store.
80
+ - Local Hugging Face adapter inference depends on a compatible environment and sufficient compute.
81
+ - This backend is a system-design layer and is not the live public Space entry point.
 
CONTRIBUTING.md ADDED
@@ -0,0 +1,62 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Contributing
2
+
3
+ Thanks for improving the Banking & Finance AI Agent. This repository is maintained as a serious AI engineering portfolio, so changes should be small, testable, and honest about what the code supports.
4
+
5
+ ## Setup
6
+
7
+ ```bash
8
+ git clone https://github.com/rakeshmadasaniai/banking-genai-portfolio.git
9
+ cd banking-genai-portfolio
10
+ python -m venv .venv
11
+ .venv\Scripts\Activate.ps1
12
+ pip install -r 01-rag-system/requirements.txt
13
+ ```
14
+
15
+ Copy the sample environment file and fill local secrets:
16
+
17
+ ```bash
18
+ copy .env.example .env
19
+ ```
20
+
21
+ Never commit real API keys or tokens.
22
+
23
+ ## Branching
24
+
25
+ - Use short descriptive branches.
26
+ - Prefer `feature/...`, `fix/...`, or `docs/...`.
27
+ - Keep deployment-risky changes separate from documentation-only changes.
28
+
29
+ ## Testing
30
+
31
+ Run the available unit tests:
32
+
33
+ ```bash
34
+ cd 01-rag-system
35
+ python -m unittest discover -s tests -p "test_*.py"
36
+ ```
37
+
38
+ When changing evaluation logic, also run the relevant scripts under `01-rag-system/evaluation`.
39
+
40
+ ## Pull Request Expectations
41
+
42
+ Each PR should include:
43
+
44
+ - What changed.
45
+ - Why it changed.
46
+ - How it was tested.
47
+ - Any limitations or follow-up work.
48
+
49
+ For AI behavior changes, include at least one before/after example and avoid unsupported accuracy claims.
50
+
51
+ ## Code Style
52
+
53
+ - Keep Python readable and typed where practical.
54
+ - Prefer small functions over large hidden control flow.
55
+ - Add comments only where the logic is non-obvious.
56
+ - Preserve existing UI design language unless the change is explicitly a redesign.
57
+
58
+ ## Documentation Style
59
+
60
+ - Use "AI agent", "grounded GenAI system", or "tool-calling AI system" when accurate.
61
+ - Avoid "AGI" or "fully autonomous" unless the implementation and evaluation clearly support that exact claim.
62
+ - Keep metrics tied to committed evaluation files.
EVALUATION.md ADDED
@@ -0,0 +1,77 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Evaluation
2
+
3
+ This document describes how the Banking & Finance AI Agent is evaluated and what should be improved next.
4
+
5
+ ## Goals
6
+
7
+ - Measure whether answers are grounded in retrieved banking context.
8
+ - Measure latency for product responsiveness.
9
+ - Compare OpenAI, Fine-Tuned, Auto, and agentic paths where configured.
10
+ - Test multilingual behavior rather than assuming English-only quality.
11
+ - Capture failure modes openly so improvements are traceable.
12
+
13
+ ## Evaluation Assets
14
+
15
+ | Asset | Location | Purpose |
16
+ |---|---|---|
17
+ | Domain query pack | `01-rag-system/evaluation/evaluation_queries.md` | Banking, AML, KYC, FDIC, RBI, Basel III, payments, and compliance prompts. |
18
+ | Multilingual query pack | `01-rag-system/evaluation/evaluation_multilingual.md` | Multilingual coverage across model modes. |
19
+ | Runner | `01-rag-system/evaluation/run_eval_sets.py` | Executes evaluation packs. |
20
+ | Summarizer | `01-rag-system/evaluation/summarize_eval_sets.py` | Produces CSV/JSON summaries. |
21
+ | Results | `01-rag-system/evaluation/results/` | Committed result snapshots. |
22
+ | Portfolio report | `01-rag-system/evaluation/reports/latest_portfolio_report.md` | Recruiter-friendly summary generated from committed artifacts. |
23
+ | Autonomy audit | `AUTONOMY_EVALUATION.md` | Honest status of the agentic/autonomous runtime. |
24
+
25
+ ## Current Snapshot
26
+
27
+ | Evaluation set | Total prompts | Available evaluated rows | Average latency | Median latency |
28
+ |---|---:|---:|---:|---:|
29
+ | Domain pack | 120 | 80 | 2037.0 ms | 2036.0 ms |
30
+ | Multilingual pack | 120 | 80 | 2031.8 ms | 2031.5 ms |
31
+
32
+ The committed export includes unavailable rows where the relevant backend was not active in the local evaluation environment. This is preserved intentionally so the results stay auditable.
33
+
34
+ ## Prompt Categories
35
+
36
+ - Banking definitions and explainers.
37
+ - AML/KYC/CDD/EDD compliance.
38
+ - FDIC deposit insurance.
39
+ - RBI and India banking compliance.
40
+ - Basel III capital and risk concepts.
41
+ - Payments and transaction-monitoring scenarios.
42
+ - Cross-jurisdiction comparison prompts.
43
+ - Multilingual banking questions.
44
+ - Agentic decision scenarios such as fraud, sanctions, short-horizon investing, and life-event planning.
45
+
46
+ ## Groundedness Scoring
47
+
48
+ The runtime uses scoring utilities to estimate:
49
+
50
+ - overlap with retrieved documents,
51
+ - completeness of the answer,
52
+ - latency quality,
53
+ - combined candidate score for routing.
54
+
55
+ These scores are useful product signals, not formal legal or regulatory validation.
56
+
57
+ ## Latency Measurement
58
+
59
+ Latency is recorded in milliseconds in model results and evaluation outputs. Current committed snapshot averages are around 2.03 seconds for available rows. Future reports should add p50, p95, p99, and per-mode latency.
60
+
61
+ ## Failure Modes To Track
62
+
63
+ - Retrieval misses for narrow regulatory facts.
64
+ - Answers that ask for clarification when enough information already exists.
65
+ - Overly generic policy-generation responses.
66
+ - Cross-jurisdiction questions that need explicit table formatting.
67
+ - Voice and upload behavior that depends on browser/runtime permissions.
68
+ - Fine-Tuned mode availability when the hosted endpoint is not configured.
69
+
70
+ ## Future Benchmark Plan
71
+
72
+ - Add a gold-answer set for 100 high-value banking and compliance questions.
73
+ - Add multilingual human review for at least five languages.
74
+ - Add separate benchmarks for retrieval-only, OpenAI, Fine-Tuned, Auto, and agentic modes.
75
+ - Add document-upload tests for PDF, DOCX, TXT, and image inputs.
76
+ - Add voice input/output smoke tests where runtime support is available.
77
+ - Track before/after scores for BM25/RRF once sparse retrieval is implemented.
README.md CHANGED
@@ -1,5 +1,5 @@
1
  ---
2
- title: Autonomous Banking & Finance AI Agent
3
  emoji: 🌎
4
  colorFrom: blue
5
  colorTo: indigo
@@ -10,290 +10,185 @@ app_file: 01-rag-system/app.py
10
  pinned: false
11
  ---
12
 
13
- # Autonomous Banking & Finance AI Agent Portfolio
14
 
15
- **Rakesh Madasani**
16
- [Live Autonomous Banking & Finance AI Agent](https://huggingface.co/spaces/RakeshMadasani/banking-finance-rag) | [Hugging Face Profile](https://huggingface.co/RakeshMadasani) | [GitHub](https://github.com/rakeshmadasaniai/banking-genai-portfolio) | [LinkedIn](https://www.linkedin.com/in/rakesh-madasani-b217b71b0/)
 
 
 
 
17
 
18
- This repository is the full build story behind my Banking & Finance Copilot: a live grounded AI product focused on USA and India banking, compliance, AML, KYC, FDIC, Basel III, and RBI workflows.
19
 
20
- I did not want this to be a one-screen chatbot demo. I wanted it to behave like a real product:
21
 
22
- - grounded on visible source material
23
- - measurable with repeatable evaluation packs
24
- - flexible across OpenAI, Fine-Tuned, and Auto routing
25
- - multilingual enough for broader banking users
26
- - strong enough to discuss as an engineering system, not just a UI
27
 
28
- ## World-Class Execution Plan
29
 
30
- I maintain a concrete week-by-week delivery plan in:
31
-
32
- - [`WORLDCLASS_WEEK1_TO_WEEK6.md`](WORLDCLASS_WEEK1_TO_WEEK6.md)
33
- - [`AUTONOMY_EVALUATION.md`](AUTONOMY_EVALUATION.md)
34
-
35
- This is the operating plan used to move the project from strong prototype quality to production-grade architecture, reliability, governance, and distribution.
36
-
37
- ## Live Product
38
-
39
- - **Live app:** [banking-finance-rag](https://huggingface.co/spaces/RakeshMadasani/banking-finance-rag)
40
- - **Dataset:** [banking-finance-qa-dataset](https://huggingface.co/datasets/RakeshMadasani/banking-finance-qa-dataset)
41
- - **Fine-tuned model:** [banking-finance-mistral-qlora](https://huggingface.co/RakeshMadasani/banking-finance-mistral-qlora)
42
-
43
- ## What This Repo Shows
44
-
45
- | Layer | What is in the repo | Why it matters |
46
- |---|---|---|
47
- | Product | A live Banking & Finance Copilot | Shows a shipped, testable AI product |
48
- | Retrieval | FAISS + banking knowledge + source cards | Keeps answers grounded and explainable |
49
- | Data | A domain-specific QA dataset | Shows data ownership, not just prompting |
50
- | Model | QLoRA fine-tuning workflow | Shows model adaptation beyond API usage |
51
- | Backend | Conversational memory API | Shows system thinking and architecture depth |
52
- | Evaluation | 120-query domain set + 120-query multilingual set | Shows repeatable measurement, not anecdotal demos |
53
-
54
- ## Where The Product Code Lives
55
-
56
- If someone lands on this repo from GitHub first, the main product code is not hidden in a separate private service. It lives directly inside [`01-rag-system`](01-rag-system):
57
-
58
- - [`01-rag-system/core`](01-rag-system/core)
59
- runtime orchestration, retrieval flow, prompts, and shared utilities
60
- - [`01-rag-system/features`](01-rag-system/features)
61
- product UI, uploads, voice output, answer cards, and interaction behavior
62
- - [`01-rag-system/models`](01-rag-system/models)
63
- OpenAI mode, Fine-Tuned mode, and Auto routing logic
64
-
65
- That structure matters because I wanted the repo to read like a real product codebase, not a single README pointing to an external demo.
66
-
67
- ## How The Portfolio Evolves
68
-
69
- This repo is one system built in layers.
70
-
71
- ### 1. Product layer
72
-
73
- I started with the user-facing assistant in [`01-rag-system`](01-rag-system). The goal was simple: if someone asks a banking or compliance question, the product should answer clearly and show the evidence behind the answer.
74
-
75
- ### 2. Data layer
76
-
77
- Once the first retrieval system worked, I created a banking QA dataset in [`02-qa-dataset`](02-qa-dataset) so the domain logic would not live only inside prompts and chunk text.
78
-
79
- ### 3. Model layer
80
-
81
- Then I fine-tuned a banking-domain adapter in [`03-qlora-finetuning`](03-qlora-finetuning) to show that I can move from application wiring into actual model adaptation.
82
-
83
- ### 4. Backend layer
84
-
85
- Finally, I added session memory and orchestration work in [`04-conversational-memory`](04-conversational-memory), which made the portfolio feel more like a real product system than a single-page demo.
86
-
87
- ## Architecture
88
 
89
- ### End-to-End System
90
 
91
- ```mermaid
92
- flowchart LR
93
- A["Curated banking knowledge"] --> B["Chunking and preprocessing"]
94
- B --> C["Embeddings"]
95
- C --> D["FAISS index"]
96
- D --> E["Shared grounded context"]
97
-
98
- A --> F["Instruction dataset generation"]
99
- F --> G["Banking QA dataset"]
100
- G --> H["QLoRA fine-tuning"]
101
- H --> I["Published banking adapter"]
102
-
103
- E --> J["OpenAI path"]
104
- E --> K["Fine-Tuned path"]
105
- E --> L["Auto routing"]
106
-
107
- J --> M["Live Banking & Finance Copilot"]
108
- K --> M
109
- L --> M
110
-
111
- M --> N["Sources, confidence, latency, uploads, read aloud"]
112
- M --> O["Conversational memory backend"]
113
- ```
114
 
115
- ### Live Product Runtime
116
 
117
  ```mermaid
118
  flowchart TD
119
- U["Question or uploaded document"] --> R["Shared retrieval"]
120
- R --> C["Grounded context"]
121
- C --> O["OpenAI mode"]
122
- C --> F["Fine-Tuned mode"]
123
- C --> A["Auto mode"]
124
-
125
- A --> S["Selection logic"]
126
- O --> X["Answer card"]
127
- F --> X
128
- S --> X
129
-
130
- X --> Y["Latency, confidence, sources, read aloud"]
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
131
  ```
132
 
133
- ## Main Product: Banking & Finance Copilot
134
-
135
- The main product lives in [`01-rag-system`](01-rag-system).
136
-
137
- What it does well:
138
- - grounded banking and compliance Q&A
139
- - OpenAI, Fine-Tuned, and Auto modes
140
- - source-backed answers with visible evidence
141
- - PDF, DOCX, TXT, and image upload support
142
- - multilingual answer support
143
- - evaluation workflows committed alongside the app
144
-
145
- ### Product Walkthrough
146
-
147
- #### A clean first impression for the product
148
-
149
- This is the opening experience of the Banking & Finance Copilot: stable sidebar, mode selector, welcome guidance, and multilingual starter questions that make the product feel usable immediately.
150
-
151
- ![Banking Copilot home experience](01-rag-system/screenshots/banking-copilot-home-experience.png)
152
-
153
- #### A grounded English answer that feels concise and useful
154
-
155
- This example shows the assistant answering a KYC question in English with a direct explanation, compact bullets, visible latency, and retrieved source support.
156
 
157
- ![English KYC answer walkthrough](01-rag-system/screenshots/english-kyc-answer-walkthrough.png)
158
 
159
- #### The same product experience working in Telugu
160
-
161
- This screenshot shows the answer structure holding up in Telugu, which is important because the product is meant to feel consistent across languages, not just translated.
162
-
163
- ![Telugu KYC answer walkthrough](01-rag-system/screenshots/telugu-kyc-answer-walkthrough.png)
164
-
165
- #### Multilingual grounding working in Chinese as well
166
-
167
- This example shows the same KYC workflow in Chinese, which helps demonstrate that multilingual support is part of the product design, not just a side feature.
168
-
169
- ![Chinese KYC answer walkthrough](01-rag-system/screenshots/chinese-kyc-answer-walkthrough.png)
170
-
171
- ## Evaluation
172
-
173
- The repo includes two larger evaluation packs inside [`01-rag-system/evaluation`](01-rag-system/evaluation):
174
-
175
- - `evaluation_queries.md`
176
- 120 domain-specific banking, AML, KYC, Basel III, FDIC, RBI, CECL, and payments questions
177
- - `evaluation_multilingual.md`
178
- 120 multilingual questions grouped across OpenAI, Fine-Tuned, and Auto modes
179
-
180
- The folder also includes:
181
-
182
- - `run_eval_sets.py` to run both evaluation packs automatically
183
- - `summarize_eval_sets.py` to summarize any generated CSV
184
- - committed result snapshots in [`01-rag-system/evaluation/results`](01-rag-system/evaluation/results)
185
- - raw CSV outputs plus JSON summaries, so the runs are inspectable rather than just summarized in prose
186
-
187
- ### Latest committed evaluation snapshots
188
 
189
- **Domain evaluation pack**
190
 
191
- | Metric | Result |
192
- |---|---|
193
- | Total prompts | 120 |
194
- | Available evaluated rows | 80 |
195
- | Average latency | 2037.0 ms |
196
- | Median latency | 2036.0 ms |
197
 
198
- **Multilingual evaluation pack**
199
 
200
  | Metric | Result |
201
  |---|---|
202
- | Total prompts | 120 |
203
- | Available evaluated rows | 80 |
204
- | Average latency | 2031.8 ms |
205
- | Median latency | 2031.5 ms |
206
-
207
- ### What those numbers mean
208
-
209
- The committed snapshot is intentionally honest about the environment it was run in:
 
210
 
211
- - the full packs contain 120 prompts each
212
- - the committed run has 80 available rows because the OpenAI path was not active in that local export
213
- - the Fine-Tuned and Auto paths still completed and produced auditable CSV/JSON artifacts
214
 
215
- I prefer showing that reality instead of pretending every backend was active in every run. A reviewer can open the raw result files, see which rows were available, and rerun the exact same packs in a fully configured environment.
216
 
217
- ### Portfolio report generator
218
 
219
- To generate a single recruiter-friendly report from committed summary artifacts:
220
 
221
- - `python 01-rag-system/evaluation/generate_portfolio_report.py`
222
- - Output: `01-rag-system/evaluation/reports/latest_portfolio_report.md`
223
 
224
- This keeps the published numbers traceable and reproducible.
225
 
226
- ### Regression tests for decision-critical behavior
227
 
228
- I also added deterministic tests for high-risk agent decisions:
229
 
230
- - `01-rag-system/tests/test_agentic_decision_engine.py`
231
 
232
- Run locally:
233
 
234
- - `cd 01-rag-system`
235
- - `python -m unittest discover -s tests -p "test_*.py"`
236
 
237
- ### Why I kept the raw result files
238
 
239
- The most valuable part of the evaluation setup is not just the summary table. It is that the repo contains:
240
 
241
- - the question packs
242
- - the runner scripts
243
- - the summarizer
244
- - the committed outputs
 
 
245
 
246
- So if someone asks, "How did you test it?" I can point to the exact prompts, exact outputs, and exact summaries rather than hand-picked screenshots.
247
 
248
- Those results matter to me because they make the product discussable in a serious way. If someone asks how I tested it, I can point to committed query packs, reproducible runners, timestamped results, and summaries instead of hand-wavy claims.
249
-
250
- ## Other Projects In The Portfolio
251
-
252
- ### [02-qa-dataset](02-qa-dataset)
253
-
254
- This is the dataset layer behind the banking system. It contains the curated QA data used to support model adaptation and domain coverage.
255
-
256
- ![Dataset screenshot](02-qa-dataset/screenshots/dataset-hf-splits.png)
257
-
258
- ### [03-qlora-finetuning](03-qlora-finetuning)
259
-
260
- This is the model adaptation layer. It shows the QLoRA workflow used to adapt a Mistral model for banking-domain answers.
261
 
262
- ![QLoRA model page screenshot](03-qlora-finetuning/screenshots/model-page-demo.png)
263
 
264
- ### [04-conversational-memory](04-conversational-memory)
 
 
 
 
 
 
 
 
265
 
266
- This is the backend layer that adds session memory, orchestration, and API structure to the broader assistant system.
267
 
268
- ## Best Entry Points In Code
 
 
 
 
 
 
269
 
270
- If someone wants to inspect the implementation rather than just the screenshots, these are the best places to start:
271
 
272
- - `01-rag-system/app.py`
273
- - `01-rag-system/core/product_runtime.py`
274
- - `01-rag-system/core/retriever.py`
275
- - `01-rag-system/features/product_ui.py`
276
- - `01-rag-system/models/auto_router.py`
277
- - `01-rag-system/models/openai_mode.py`
278
- - `01-rag-system/models/finetuned_mode.py`
279
- - `01-rag-system/evaluation/run_eval_sets.py`
280
- - `01-rag-system/evaluation/summarize_eval_sets.py`
281
- - `02-qa-dataset/generate_dataset.py`
282
- - `03-qlora-finetuning/inference_demo.py`
283
- - `04-conversational-memory/app/main.py`
284
- - `04-conversational-memory/app/rag_chain.py`
285
 
286
- ## Repo Structure
287
 
288
- ```text
289
- banking-genai-portfolio/
290
- |-- README.md
291
- |-- 01-rag-system/
292
- |-- 02-qa-dataset/
293
- |-- 03-qlora-finetuning/
294
- `-- 04-conversational-memory/
295
- ```
296
 
297
- ## Closing Note
298
 
299
- The part I value most in this portfolio is not that it calls an LLM. It is that the repo shows the full path from idea to product: retrieval, data, model work, backend orchestration, evaluation, and a live user-facing deployment that someone can test today.
 
1
  ---
2
+ title: Banking & Finance AI Agent
3
  emoji: 🌎
4
  colorFrom: blue
5
  colorTo: indigo
 
10
  pinned: false
11
  ---
12
 
13
+ # Banking & Finance AI Agent
14
 
15
+ ![Python](https://img.shields.io/badge/Python-3.10-blue)
16
+ ![Hugging Face Spaces](https://img.shields.io/badge/Hugging%20Face-Spaces-yellow)
17
+ ![Streamlit](https://img.shields.io/badge/Streamlit-1.56-red)
18
+ ![FastAPI](https://img.shields.io/badge/FastAPI-memory%20backend-009688)
19
+ ![License](https://img.shields.io/badge/License-MIT-green)
20
+ ![Status](https://img.shields.io/badge/Status-Active-brightgreen)
21
 
22
+ **Production-grade GenAI system for grounded banking, compliance, and financial knowledge workflows.**
23
 
24
+ I built this system to answer one question: what does it take to move a GenAI product beyond a chatbot demo and into a reliable, measurable AI system?
25
 
26
+ The result is a live Banking & Finance AI Agent with retrieval, model routing, domain adaptation, conversational memory work, evaluation packs, multilingual UX, upload workflows, voice support, and an autonomy audit. It is intentionally positioned as an AI agent and grounded GenAI platform, not AGI and not an overclaimed fully autonomous production system.
 
 
 
 
27
 
28
+ ## Live Links
29
 
30
+ | Asset | Link |
31
+ |---|---|
32
+ | Live app | [Hugging Face Space](https://huggingface.co/spaces/RakeshMadasani/banking-finance-rag) |
33
+ | GitHub repository | [banking-genai-portfolio](https://github.com/rakeshmadasaniai/banking-genai-portfolio) |
34
+ | Fine-tuned model | [banking-finance-mistral-qlora](https://huggingface.co/RakeshMadasani/banking-finance-mistral-qlora) |
35
+ | Dataset | [banking-finance-qa-dataset](https://huggingface.co/datasets/RakeshMadasani/banking-finance-qa-dataset) |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
36
 
37
+ ## What This System Does
38
 
39
+ - Answers banking, finance, AML, KYC, FDIC, Basel III, RBI, and compliance questions with retrieved context.
40
+ - Supports OpenAI, Fine-Tuned, Auto, Agentic Workspace, and Autonomous Max paths where configured.
41
+ - Renders source-grounded answer cards with latency, confidence, retrieved chunks, source cards, copy/export actions, and read-aloud controls.
42
+ - Accepts text, document uploads, image-supported workflows, multilingual prompts, and voice input/output paths.
43
+ - Includes a published BankingQA-3K dataset and a QLoRA Mistral-7B adapter for domain model adaptation.
44
+ - Ships repeatable evaluation packs, committed result snapshots, and an autonomy evaluation note instead of only screenshots.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
45
 
46
+ ## Architecture Overview
47
 
48
  ```mermaid
49
  flowchart TD
50
+ U["User input: text, document, image, or voice"] --> IR["Input router"]
51
+ IR --> UP["Upload and document parsing"]
52
+ IR --> VI["Voice input path"]
53
+ IR --> Q["Normalized user query"]
54
+
55
+ UP --> KB["Runtime knowledge context"]
56
+ Q --> RC["Retrieval coordinator"]
57
+ KB --> RC
58
+
59
+ RC --> FAISS["Implemented: FAISS dense vector search"]
60
+ RC -. "roadmap" .-> BM25["Planned: BM25 sparse search"]
61
+ BM25 -. "roadmap" .-> RRF["Planned: reciprocal rank fusion"]
62
+ FAISS --> GC["Grounded context"]
63
+ RRF -. "future hybrid context" .-> GC
64
+
65
+ GC --> ORCH["LLM orchestration layer"]
66
+ ORCH --> OAI["OpenAI mode"]
67
+ ORCH --> FT["Fine-Tuned mode"]
68
+ ORCH --> AUTO["Auto routing"]
69
+ ORCH --> AGENT["Agentic / Autonomous modes"]
70
+
71
+ OAI --> EVAL["Evaluation + confidence scoring"]
72
+ FT --> EVAL
73
+ AUTO --> EVAL
74
+ AGENT --> EVAL
75
+
76
+ EVAL --> RESP["Response with sources, confidence, latency, and actions"]
77
+ RESP --> MEM["Session memory and audit context"]
78
+ MEM --> ORCH
79
  ```
80
 
81
+ ## System Design
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
82
 
83
+ The system is built in four layers:
84
 
85
+ | Layer | Runtime area | Purpose |
86
+ |---|---|---|
87
+ | AI agent runtime | `01-rag-system` | Live Streamlit product, retrieval, orchestration, source-grounded UI, uploads, voice, and agent paths. |
88
+ | BankingQA dataset | `02-qa-dataset` | 3,002-pair instruction dataset covering banking, compliance, AML, KYC, Basel III, FDIC, RBI, and finance topics. |
89
+ | Domain model adaptation | `03-qlora-finetuning` | QLoRA workflow for adapting Mistral-7B-Instruct-v0.3 to banking and financial compliance terminology. |
90
+ | Memory and orchestration | `04-conversational-memory` | FastAPI memory backend with session handling, history management, summarization, and backend comparison. |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
91
 
92
+ The current live Space keeps the original folder path `01-rag-system` so existing Hugging Face deployment URLs and `app_file` metadata continue to work. Documentation now positions that folder as the AI agent runtime while preserving compatibility.
93
 
94
+ ## Evaluation Metrics
 
 
 
 
 
95
 
96
+ These numbers are from committed project files and should be read as traceable project evidence, not marketing claims.
97
 
98
  | Metric | Result |
99
  |---|---|
100
+ | Evaluation prompts | 240 total prompts across domain and multilingual packs |
101
+ | Latest available evaluated rows | 160 rows across committed domain and multilingual snapshots |
102
+ | Domain pack average latency | 2037.0 ms |
103
+ | Multilingual pack average latency | 2031.8 ms |
104
+ | Dataset size | 3,002 QA pairs |
105
+ | Fine-tuning method | Mistral-7B QLoRA |
106
+ | Final train loss | 1.13 |
107
+ | Current retrieval implementation | FAISS dense vector retrieval |
108
+ | Sparse retrieval / RRF | Roadmap item, not claimed as live implementation |
109
 
110
+ For the current autonomy positioning, see [`AUTONOMY_EVALUATION.md`](AUTONOMY_EVALUATION.md). For the generated portfolio report, see [`01-rag-system/evaluation/reports/latest_portfolio_report.md`](01-rag-system/evaluation/reports/latest_portfolio_report.md).
 
 
111
 
112
+ ## Key Capabilities
113
 
114
+ ### Grounded Banking Answers
115
 
116
+ The runtime retrieves banking material before generation, then presents answers with source cards, confidence labels, latency, and chunk metadata.
117
 
118
+ ### Model Orchestration
 
119
 
120
+ OpenAI mode provides a stable general path, Fine-Tuned mode connects the domain adapter path where hosted inference is configured, and Auto mode scores candidate answers based on groundedness, completeness, and latency.
121
 
122
+ ### Agentic Runtime Work
123
 
124
+ The repo includes agentic/autonomous runtime work with tool-style execution traces and autonomy evaluation. This is presented honestly as a tool-calling AI system and agentic workflow layer, not as AGI.
125
 
126
+ ### Domain Data and Model Adaptation
127
 
128
+ The dataset and QLoRA adapter show the system is not only prompt engineering. It includes a reusable data asset and a domain-adapted model artifact.
129
 
130
+ ### Memory and API Layer
 
131
 
132
+ The FastAPI memory backend demonstrates session-aware conversation handling, summarization/truncation, health checks, and backend comparison endpoints.
133
 
134
+ ## Demo Workflow
135
 
136
+ 1. Open the [live Space](https://huggingface.co/spaces/RakeshMadasani/banking-finance-rag).
137
+ 2. Ask a banking or compliance question such as `What are the main KYC requirements for banks?`.
138
+ 3. Switch modes to compare OpenAI, Fine-Tuned, Auto, and agentic paths where configured.
139
+ 4. Upload a PDF, DOCX, or TXT document and ask a document-grounded question.
140
+ 5. Inspect confidence, source cards, latency, retrieved chunks, and read-aloud output.
141
+ 6. Review evaluation artifacts under `01-rag-system/evaluation`.
142
 
143
+ ## Repository Structure
144
 
145
+ ```text
146
+ banking-genai-portfolio/
147
+ |-- README.md
148
+ |-- ROADMAP.md
149
+ |-- EVALUATION.md
150
+ |-- SYSTEM_DESIGN.md
151
+ |-- CONTRIBUTING.md
152
+ |-- AUTONOMY_EVALUATION.md
153
+ |-- 01-rag-system/ # AI agent runtime and live Streamlit app
154
+ |-- 02-qa-dataset/ # BankingQA-3K dataset build/publish workflow
155
+ |-- 03-qlora-finetuning/ # Mistral-7B QLoRA adaptation workflow
156
+ `-- 04-conversational-memory/ # FastAPI memory and orchestration backend
157
+ ```
158
 
159
+ ## Run Locally
160
 
161
+ ```bash
162
+ git clone https://github.com/rakeshmadasaniai/banking-genai-portfolio.git
163
+ cd banking-genai-portfolio
164
+ python -m venv .venv
165
+ .venv\Scripts\Activate.ps1
166
+ pip install -r 01-rag-system/requirements.txt
167
+ copy .env.example .env
168
+ streamlit run 01-rag-system/app.py
169
+ ```
170
 
171
+ Minimum environment:
172
 
173
+ ```env
174
+ OPENAI_API_KEY=
175
+ HF_TOKEN=
176
+ MODEL_MODE=OpenAI
177
+ TOP_K=4
178
+ TEMPERATURE=0.2
179
+ ```
180
 
181
+ ## Known Limitations
182
 
183
+ - The live retrieval path is FAISS dense search; BM25 and reciprocal rank fusion are documented as roadmap work until implemented in code.
184
+ - Fine-Tuned mode depends on an available hosted endpoint or compatible local inference environment.
185
+ - Streamlit session state is not durable across browser restarts; the separate FastAPI memory backend demonstrates the production direction.
186
+ - The system is educational and portfolio-grade; it is not legal, financial, investment, or compliance advice.
 
 
 
 
 
 
 
 
 
187
 
188
+ ## Roadmap
189
 
190
+ See [`ROADMAP.md`](ROADMAP.md) for the planned reliability, agentic architecture, governance, and production-hardening phases.
 
 
 
 
 
 
 
191
 
192
+ ## License and Contact
193
 
194
+ This repository is intended as an AI engineering portfolio and educational system. For questions, reach out through [GitHub](https://github.com/rakeshmadasaniai), [Hugging Face](https://huggingface.co/RakeshMadasani), or [LinkedIn](https://www.linkedin.com/in/rakesh-madasani-b217b71b0/).
ROADMAP.md ADDED
@@ -0,0 +1,49 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Roadmap
2
+
3
+ This roadmap keeps the project moving toward a stronger production-grade Banking & Finance AI Agent without overclaiming what is already implemented.
4
+
5
+ ## Phase 1 - Reliability
6
+
7
+ - Fill any missing evaluation tables with reproducible result snapshots.
8
+ - Add regression tests for high-risk banking, compliance, investment, and multilingual paths.
9
+ - Improve source-grounding checks and weak-retrieval fallbacks.
10
+ - Track p50, p95, and p99 latency instead of only averages.
11
+ - Add deterministic smoke tests for the Streamlit runtime startup path.
12
+
13
+ ## Phase 2 - Agentic Architecture
14
+
15
+ - Formalize a Planner -> Executor -> Verifier loop.
16
+ - Standardize tool-use traces across all agentic modes.
17
+ - Add retry policy, max-step budgets, and clear stop conditions.
18
+ - Separate "answer directly" from "act with tools" using an explicit decision layer.
19
+ - Add durable task-state persistence for long-running autonomous workflows.
20
+
21
+ ## Phase 3 - Governance
22
+
23
+ - Add compliance guardrails for AML, KYC, sanctions, investment-risk, and crisis scenarios.
24
+ - Add structured audit logs for tool calls, verification decisions, and fallback paths.
25
+ - Add escalation policy for unsupported legal, compliance, or investment-advice requests.
26
+ - Add risk scoring for answers that combine regulated finance and user-specific facts.
27
+ - Create reviewer-friendly model cards and dataset cards with limitations clearly stated.
28
+
29
+ ## Phase 4 - Production Hardening
30
+
31
+ - Add Docker deployment for the Streamlit app and FastAPI memory backend.
32
+ - Add CI/CD for tests, linting, and evaluation smoke checks.
33
+ - Add API documentation for the memory backend.
34
+ - Add monitoring for latency, tool failure rates, retrieval misses, and user-facing errors.
35
+ - Add versioned releases and changelogs.
36
+
37
+ ## Phase 5 - Retrieval Upgrade
38
+
39
+ - Add BM25 sparse retrieval.
40
+ - Add reciprocal rank fusion between FAISS dense retrieval and BM25 sparse retrieval.
41
+ - Add retrieval ablation tests to quantify dense-only versus hybrid retrieval quality.
42
+ - Add query rewriting for difficult regulatory and multilingual prompts.
43
+
44
+ ## Phase 6 - Portfolio Distribution
45
+
46
+ - Publish a concise technical write-up explaining the system design and evaluation.
47
+ - Mirror the live Hugging Face Space state into GitHub branches or releases.
48
+ - Add demo clips or GIFs that show source cards, agent trace, uploads, and voice output.
49
+ - Keep README claims tied to committed code, result files, and reproducible scripts.
SYSTEM_DESIGN.md ADDED
@@ -0,0 +1,118 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # System Design
2
+
3
+ ## Problem
4
+
5
+ Banking and finance users need answers that are clear, grounded, and measurable. A normal chatbot can produce fluent text, but regulated domains require source visibility, confidence signals, evaluation, and careful handling of user-specific finance or compliance scenarios.
6
+
7
+ ## Goals
8
+
9
+ - Provide a live AI workflow interface for banking and compliance questions.
10
+ - Ground answers in curated banking material and uploaded documents.
11
+ - Support multiple model paths and routing decisions.
12
+ - Track latency, confidence, source usage, and evaluation outputs.
13
+ - Demonstrate the full system chain: data, retrieval, fine-tuning, memory, evaluation, and deployment.
14
+
15
+ ## Non-Goals
16
+
17
+ - Do not claim AGI.
18
+ - Do not claim formal legal, investment, financial, or compliance advice.
19
+ - Do not claim BM25/RRF as implemented until the code path exists.
20
+ - Do not claim fully autonomous production operation without durable planning, monitoring, governance, and external action controls.
21
+
22
+ ## Architecture
23
+
24
+ ```mermaid
25
+ flowchart TD
26
+ A["User input"] --> B["Input router"]
27
+ B --> C["Text / upload / voice handling"]
28
+ C --> D["Chunking + embeddings"]
29
+ D --> E["FAISS dense retrieval"]
30
+ E --> F["Grounded context"]
31
+ F --> G["Model orchestration"]
32
+ G --> H["OpenAI mode"]
33
+ G --> I["Fine-Tuned mode"]
34
+ G --> J["Auto mode"]
35
+ G --> K["Agentic / Autonomous modes"]
36
+ H --> L["Confidence + source rendering"]
37
+ I --> L
38
+ J --> L
39
+ K --> L
40
+ L --> M["Response UI"]
41
+ M --> N["Session memory / audit context"]
42
+ ```
43
+
44
+ ## Retrieval Design
45
+
46
+ Current implementation:
47
+
48
+ - Documents are chunked and embedded.
49
+ - FAISS provides dense vector search.
50
+ - Retrieved context is passed into answer generation and source cards.
51
+
52
+ Planned retrieval hardening:
53
+
54
+ - Add BM25 sparse retrieval for exact regulatory terms.
55
+ - Fuse dense and sparse results with reciprocal rank fusion.
56
+ - Add retrieval evaluation comparing dense-only and hybrid retrieval.
57
+
58
+ ## Model Orchestration
59
+
60
+ The runtime supports several paths:
61
+
62
+ - OpenAI mode for stable general-purpose answers.
63
+ - Fine-Tuned mode for the banking-domain adapter path.
64
+ - Auto mode for scoring candidate answers and selecting a winner.
65
+ - Agentic/autonomous modes for tool-style workflows and execution traces where configured.
66
+
67
+ Candidate scoring uses groundedness, completeness, and latency signals. This makes routing inspectable instead of hidden.
68
+
69
+ ## Memory Design
70
+
71
+ The live Streamlit runtime uses session state for chat/session continuity during a browser session. The separate FastAPI memory backend demonstrates the production direction:
72
+
73
+ - session IDs,
74
+ - retained recent turns,
75
+ - summarization/truncation,
76
+ - comparison endpoints,
77
+ - health checks.
78
+
79
+ Production memory should move to durable storage such as Redis or Postgres.
80
+
81
+ ## Evaluation Design
82
+
83
+ Evaluation is repository-native:
84
+
85
+ - domain prompt packs,
86
+ - multilingual prompt packs,
87
+ - runner scripts,
88
+ - summarizer scripts,
89
+ - committed CSV/JSON outputs,
90
+ - generated portfolio report,
91
+ - autonomy audit.
92
+
93
+ This structure lets reviewers inspect prompts and outputs rather than relying on hand-picked examples.
94
+
95
+ ## Reliability Considerations
96
+
97
+ - Weak retrieval should be surfaced clearly.
98
+ - Mode failures should degrade gracefully.
99
+ - Fine-Tuned endpoint unavailability should not crash the app.
100
+ - Upload parsing should handle unsupported or malformed files safely.
101
+ - Agentic workflows need max-step limits, retry budgets, and stop conditions.
102
+ - Regulated finance answers should include disclaimers and avoid unsupported personalized advice.
103
+
104
+ ## Tradeoffs
105
+
106
+ - Streamlit is excellent for fast product iteration, but not a full production backend by itself.
107
+ - FAISS dense retrieval is simple and fast, but sparse retrieval is needed for exact regulatory threshold matching.
108
+ - QLoRA adapters are efficient for domain adaptation, but inference still requires a compatible base model environment.
109
+ - Agentic behavior improves complex workflows, but it adds latency and requires strict controls.
110
+
111
+ ## Future Improvements
112
+
113
+ - Implement BM25 + reciprocal rank fusion.
114
+ - Add durable memory and audit logs.
115
+ - Add Docker and CI/CD.
116
+ - Add p95/p99 latency monitoring.
117
+ - Add benchmarked document-upload and voice tests.
118
+ - Add governance guardrails for high-risk compliance and investment scenarios.
requirements.txt CHANGED
@@ -1,12 +1,16 @@
 
1
  langchain==0.1.20
2
  langchain-community==0.0.38
3
  langchain-openai==0.1.6
4
- openai==1.30.0
 
5
  httpx<0.28
6
  faiss-cpu
7
  sentence-transformers==2.7.0
8
  pypdf==4.2.0
 
9
  python-dotenv==1.0.1
10
  huggingface-hub==0.23.0
11
  python-docx==1.1.2
12
  streamlit-mic-recorder==0.0.8
 
 
1
+ streamlit==1.56.0
2
  langchain==0.1.20
3
  langchain-community==0.0.38
4
  langchain-openai==0.1.6
5
+ openai>=1.40.0
6
+ pydantic>=2.0.0
7
  httpx<0.28
8
  faiss-cpu
9
  sentence-transformers==2.7.0
10
  pypdf==4.2.0
11
+ PyPDF2>=3.0.0
12
  python-dotenv==1.0.1
13
  huggingface-hub==0.23.0
14
  python-docx==1.1.2
15
  streamlit-mic-recorder==0.0.8
16
+ yfinance