File size: 7,107 Bytes
346f08f
6ba3ef3
 
 
 
 
b534bcc
 
6ba3ef3
346f08f
 
 
6ba3ef3
346f08f
6ba3ef3
 
 
 
 
346f08f
6ba3ef3
 
4eda25e
6ba3ef3
 
 
 
 
 
 
48c9780
 
 
 
 
 
 
 
6ba3ef3
 
 
 
47541b7
 
6ba3ef3
 
 
 
 
 
 
8db8a8c
 
 
6ba3ef3
 
8db8a8c
 
 
6ba3ef3
 
 
 
 
 
 
 
 
8db8a8c
47541b7
6ba3ef3
6f598ab
 
 
 
47541b7
6ba3ef3
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
47541b7
6ba3ef3
47541b7
 
6ba3ef3
 
 
 
 
 
47541b7
 
 
 
 
 
087a312
47541b7
 
 
 
087a312
47541b7
 
 
 
 
087a312
 
47541b7
087a312
 
 
 
47541b7
087a312
 
 
 
 
 
6ba3ef3
 
 
 
 
47541b7
087a312
47541b7
 
6ba3ef3
 
 
 
 
087a312
 
 
 
ca49d1a
6ba3ef3
 
 
 
 
 
 
 
 
 
47541b7
 
6ba3ef3
 
47541b7
 
 
 
6ba3ef3
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
---
title: FinChat
emoji: πŸ’¬
colorFrom: indigo
colorTo: blue
sdk: streamlit
sdk_version: 1.58.0
python_version: "3.12"
app_file: app.py
pinned: false
---

# πŸ’¬ FinChat β€” Chat with SEC 10-K Filings

FinChat is a **Retrieval-Augmented Generation (RAG)** chatbot that answers
questions about public companies using their **SEC 10-K annual filings**.
Ask *"What are AMD's main business risks?"* and FinChat finds the relevant
passages in the filings and answers β€” **grounded in the source, with
citations** β€” instead of making things up.

> Portfolio project Β· Retrieval-Augmented Generation over financial documents.

**πŸ”— Live demo:** <https://huggingface.co/spaces/dahutapea/Finchat>

---

## ✨ Features

- **Grounded answers with citations** β€” every response is backed by excerpts
  from real 10-K filings, shown in an expandable *Sources* panel.
- **Answers financial figures, ratios & trends (hybrid RAG)** β€” plain text RAG
  can't read numbers out of financial-statement tables. FinChat extracts each
  filing's **XBRL** structured financials, **computes standard ratios**
  (margins, liquidity, returns, EBITDA, turnover, free cash flow)
  deterministically in Python, and builds **year-over-year trend** facts β€” then
  a hybrid retriever *guarantees* these are in context for numeric questions.
  So *"Apple's FY2023 revenue?"* β†’ **$383.29B**, *"quick ratio?"* β†’ **0.94**,
  *"did its margin improve YoY?"* β†’ answered straight from the data.
- **Query routing ("knows where to look")** β€” FinChat detects which company a
  question is about and searches *only* that company's filings via metadata
  filtering, with graceful semantic fallback when the company is ambiguous.
- **Refuses to hallucinate** β€” if the answer isn't in the filings, it says so.
- **Benchmarked** β€” evaluated by an LLM-as-judge on a capability gold set
  (100%) and the external FinanceBench benchmark (see [Evaluation](#-evaluation)).
- **100% free stack** β€” local embeddings + a free LLM API. No paid keys.

---

## πŸ—οΈ Architecture

```
INGESTION (once) β€” two tracks per filing
  10-K TEXT       ──► split into chunks ───────────┐
  XBRL FINANCIALS ──► "label: value" fact chunks ──┴─► embed ──► ChromaDB

QUERY (per question)
  question ──► detect company ──► HYBRID retrieve
        (semantic text chunks + guaranteed XBRL statements for numeric Qs)
        ──► LLM ──► grounded answer + citations
```

| Layer         | Tool                                              |
|---------------|---------------------------------------------------|
| Orchestration | LangChain                                         |
| Embeddings    | `BAAI/bge-small-en-v1.5` (local, free)            |
| Vector store  | ChromaDB (persisted locally)                      |
| LLM           | Llama 3.3 70B via Groq (free)                     |
| UI            | Streamlit                                         |
| Data          | SEC 10-K text **+ XBRL financials** via `edgartools` |
| Evaluation    | Capability gold set + FinanceBench, LLM-as-judge  |

**Corpus β€” 25 recognizable companies (FY2021–2024 10-Ks):** Apple, Microsoft,
Alphabet (Google), Amazon, NVIDIA, Tesla, AMD, JPMorgan Chase, American
Express, Boeing, Walmart, PepsiCo, Coca-Cola, Amcor, 3M, Johnson & Johnson,
CVS Health, Pfizer, AES, Verizon, Best Buy, Adobe, Ulta Beauty, Nike, and
Corning. Edit the list in [`src/config.py`](src/config.py).

---

## πŸš€ Setup

```bash
# 1. Create & activate a virtual environment
python -m venv venv
venv\Scripts\activate          # Windows
# source venv/bin/activate     # macOS / Linux

# 2. Install dependencies
pip install -r requirements.txt

# 3. Add your Groq API key: copy .env.example to .env and paste your key
#    Get a free key at https://console.groq.com
```

## πŸ› οΈ Usage

```bash
python -m src.ingest            # fetch filings from EDGAR + build the index (once)
streamlit run app.py            # launch the chatbot
python -m eval.run_gold         # capability eval (qualitative Q&A)
python -m eval.run_eval         # FinanceBench eval
```

---

## πŸ“Š Evaluation

FinChat is graded by an **LLM-as-judge** two ways: on the task it's built for,
and against a hard external benchmark.

**1. Capability β€” qualitative document Q&A**
([`eval/gold_results.md`](eval/gold_results.md))

A 15-question gold set (business, segments, products) across the corpus, with
reference answers from the filings.

| CORRECT | PARTIAL | INCORRECT | Accuracy |
|---|---|---|---|
| 13 | 2 | 0 | **93%** |

**2. FinanceBench β€” hard external benchmark**
([`eval/results.md`](eval/results.md))

Scored on [FinanceBench](https://huggingface.co/datasets/PatronusAI/financebench)
questions whose company + fiscal year is in the corpus. Adding the **hybrid XBRL
financials + computed-ratios layer** more than **doubled** the score:

| Setup | Accuracy | metrics-generated | domain-relevant |
|---|---|---|---|
| Text-only RAG | 20% (6/30) | 0% | 24% |
| **+ XBRL financials & ratios** | **45%** (13.5/30) | **50%** | **50%** |

The jump comes from numeric questions the text-only system couldn't touch β€”
quick ratio, gross-margin change, inventory turnover, working capital, dividend
payout β€” now answered from structured data. The remaining gap is **multi-step
reasoning** (*"excluding M&A, which segment dragged margins?"*), which needs
deeper analytical logic (future work). For context, GPT-4 in a naive RAG setup
scores **~19%** on FinanceBench.

---

## ☁️ Deployment

Deployed to Hugging Face Spaces (free) β€” see **[DEPLOY.md](DEPLOY.md)**. The
~21k-chunk vector store is prebuilt and shipped with the repo via **git-lfs**,
so the Space starts instantly with no rebuild; `config.py` auto-detects the
committed index.

---

## ⚠️ Limitations

- **Financial figures and standard ratios** (margins, liquidity, returns, FCF)
  are answered from XBRL data + deterministic computation. **Multi-step
  analytical reasoning** (e.g. segment-level margin attribution) is the
  remaining gap.
- The corpus is scoped to 25 companies' recent 10-Ks to stay laptop-friendly.
- Not financial advice β€” a portfolio/educational project.

---

## πŸ“ Project structure

```
.
β”œβ”€β”€ app.py                 # Streamlit chat UI
β”œβ”€β”€ src/
β”‚   β”œβ”€β”€ config.py          # all tunable settings (target companies, models)
β”‚   β”œβ”€β”€ ingest.py          # fetch 10-Ks from EDGAR β†’ chunk β†’ embed β†’ store
β”‚   └── rag.py             # retrieval + generation + query routing
β”œβ”€β”€ eval/
β”‚   β”œβ”€β”€ gold_set.py        # capability questions + reference answers
β”‚   β”œβ”€β”€ run_gold.py        # capability eval (qualitative Q&A)
β”‚   β”œβ”€β”€ run_eval.py        # FinanceBench eval harness
β”‚   └── *_results.md       # evaluation reports
β”œβ”€β”€ .streamlit/config.toml # Streamlit settings
β”œβ”€β”€ requirements.txt
β”œβ”€β”€ .env.example           # template for your API key
β”œβ”€β”€ DEPLOY.md              # Hugging Face Spaces deploy guide
└── README.md
```