File size: 7,107 Bytes
346f08f 6ba3ef3 b534bcc 6ba3ef3 346f08f 6ba3ef3 346f08f 6ba3ef3 346f08f 6ba3ef3 4eda25e 6ba3ef3 48c9780 6ba3ef3 47541b7 6ba3ef3 8db8a8c 6ba3ef3 8db8a8c 6ba3ef3 8db8a8c 47541b7 6ba3ef3 6f598ab 47541b7 6ba3ef3 47541b7 6ba3ef3 47541b7 6ba3ef3 47541b7 087a312 47541b7 087a312 47541b7 087a312 47541b7 087a312 47541b7 087a312 6ba3ef3 47541b7 087a312 47541b7 6ba3ef3 087a312 ca49d1a 6ba3ef3 47541b7 6ba3ef3 47541b7 6ba3ef3 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 | ---
title: FinChat
emoji: π¬
colorFrom: indigo
colorTo: blue
sdk: streamlit
sdk_version: 1.58.0
python_version: "3.12"
app_file: app.py
pinned: false
---
# π¬ FinChat β Chat with SEC 10-K Filings
FinChat is a **Retrieval-Augmented Generation (RAG)** chatbot that answers
questions about public companies using their **SEC 10-K annual filings**.
Ask *"What are AMD's main business risks?"* and FinChat finds the relevant
passages in the filings and answers β **grounded in the source, with
citations** β instead of making things up.
> Portfolio project Β· Retrieval-Augmented Generation over financial documents.
**π Live demo:** <https://huggingface.co/spaces/dahutapea/Finchat>
---
## β¨ Features
- **Grounded answers with citations** β every response is backed by excerpts
from real 10-K filings, shown in an expandable *Sources* panel.
- **Answers financial figures, ratios & trends (hybrid RAG)** β plain text RAG
can't read numbers out of financial-statement tables. FinChat extracts each
filing's **XBRL** structured financials, **computes standard ratios**
(margins, liquidity, returns, EBITDA, turnover, free cash flow)
deterministically in Python, and builds **year-over-year trend** facts β then
a hybrid retriever *guarantees* these are in context for numeric questions.
So *"Apple's FY2023 revenue?"* β **$383.29B**, *"quick ratio?"* β **0.94**,
*"did its margin improve YoY?"* β answered straight from the data.
- **Query routing ("knows where to look")** β FinChat detects which company a
question is about and searches *only* that company's filings via metadata
filtering, with graceful semantic fallback when the company is ambiguous.
- **Refuses to hallucinate** β if the answer isn't in the filings, it says so.
- **Benchmarked** β evaluated by an LLM-as-judge on a capability gold set
(100%) and the external FinanceBench benchmark (see [Evaluation](#-evaluation)).
- **100% free stack** β local embeddings + a free LLM API. No paid keys.
---
## ποΈ Architecture
```
INGESTION (once) β two tracks per filing
10-K TEXT βββΊ split into chunks ββββββββββββ
XBRL FINANCIALS βββΊ "label: value" fact chunks βββ΄ββΊ embed βββΊ ChromaDB
QUERY (per question)
question βββΊ detect company βββΊ HYBRID retrieve
(semantic text chunks + guaranteed XBRL statements for numeric Qs)
βββΊ LLM βββΊ grounded answer + citations
```
| Layer | Tool |
|---------------|---------------------------------------------------|
| Orchestration | LangChain |
| Embeddings | `BAAI/bge-small-en-v1.5` (local, free) |
| Vector store | ChromaDB (persisted locally) |
| LLM | Llama 3.3 70B via Groq (free) |
| UI | Streamlit |
| Data | SEC 10-K text **+ XBRL financials** via `edgartools` |
| Evaluation | Capability gold set + FinanceBench, LLM-as-judge |
**Corpus β 25 recognizable companies (FY2021β2024 10-Ks):** Apple, Microsoft,
Alphabet (Google), Amazon, NVIDIA, Tesla, AMD, JPMorgan Chase, American
Express, Boeing, Walmart, PepsiCo, Coca-Cola, Amcor, 3M, Johnson & Johnson,
CVS Health, Pfizer, AES, Verizon, Best Buy, Adobe, Ulta Beauty, Nike, and
Corning. Edit the list in [`src/config.py`](src/config.py).
---
## π Setup
```bash
# 1. Create & activate a virtual environment
python -m venv venv
venv\Scripts\activate # Windows
# source venv/bin/activate # macOS / Linux
# 2. Install dependencies
pip install -r requirements.txt
# 3. Add your Groq API key: copy .env.example to .env and paste your key
# Get a free key at https://console.groq.com
```
## π οΈ Usage
```bash
python -m src.ingest # fetch filings from EDGAR + build the index (once)
streamlit run app.py # launch the chatbot
python -m eval.run_gold # capability eval (qualitative Q&A)
python -m eval.run_eval # FinanceBench eval
```
---
## π Evaluation
FinChat is graded by an **LLM-as-judge** two ways: on the task it's built for,
and against a hard external benchmark.
**1. Capability β qualitative document Q&A**
([`eval/gold_results.md`](eval/gold_results.md))
A 15-question gold set (business, segments, products) across the corpus, with
reference answers from the filings.
| CORRECT | PARTIAL | INCORRECT | Accuracy |
|---|---|---|---|
| 13 | 2 | 0 | **93%** |
**2. FinanceBench β hard external benchmark**
([`eval/results.md`](eval/results.md))
Scored on [FinanceBench](https://huggingface.co/datasets/PatronusAI/financebench)
questions whose company + fiscal year is in the corpus. Adding the **hybrid XBRL
financials + computed-ratios layer** more than **doubled** the score:
| Setup | Accuracy | metrics-generated | domain-relevant |
|---|---|---|---|
| Text-only RAG | 20% (6/30) | 0% | 24% |
| **+ XBRL financials & ratios** | **45%** (13.5/30) | **50%** | **50%** |
The jump comes from numeric questions the text-only system couldn't touch β
quick ratio, gross-margin change, inventory turnover, working capital, dividend
payout β now answered from structured data. The remaining gap is **multi-step
reasoning** (*"excluding M&A, which segment dragged margins?"*), which needs
deeper analytical logic (future work). For context, GPT-4 in a naive RAG setup
scores **~19%** on FinanceBench.
---
## βοΈ Deployment
Deployed to Hugging Face Spaces (free) β see **[DEPLOY.md](DEPLOY.md)**. The
~21k-chunk vector store is prebuilt and shipped with the repo via **git-lfs**,
so the Space starts instantly with no rebuild; `config.py` auto-detects the
committed index.
---
## β οΈ Limitations
- **Financial figures and standard ratios** (margins, liquidity, returns, FCF)
are answered from XBRL data + deterministic computation. **Multi-step
analytical reasoning** (e.g. segment-level margin attribution) is the
remaining gap.
- The corpus is scoped to 25 companies' recent 10-Ks to stay laptop-friendly.
- Not financial advice β a portfolio/educational project.
---
## π Project structure
```
.
βββ app.py # Streamlit chat UI
βββ src/
β βββ config.py # all tunable settings (target companies, models)
β βββ ingest.py # fetch 10-Ks from EDGAR β chunk β embed β store
β βββ rag.py # retrieval + generation + query routing
βββ eval/
β βββ gold_set.py # capability questions + reference answers
β βββ run_gold.py # capability eval (qualitative Q&A)
β βββ run_eval.py # FinanceBench eval harness
β βββ *_results.md # evaluation reports
βββ .streamlit/config.toml # Streamlit settings
βββ requirements.txt
βββ .env.example # template for your API key
βββ DEPLOY.md # Hugging Face Spaces deploy guide
βββ README.md
```
|