dahutapea commited on
Commit
8db8a8c
·
1 Parent(s): ca49d1a

Add hybrid RAG: XBRL structured-financials layer for numeric questions

Browse files
README.md CHANGED
@@ -28,6 +28,12 @@ citations** — instead of making things up.
28
 
29
  - **Grounded answers with citations** — every response is backed by excerpts
30
  from real 10-K filings, shown in an expandable *Sources* panel.
 
 
 
 
 
 
31
  - **Query routing ("knows where to look")** — FinChat detects which company a
32
  question is about and searches *only* that company's filings via metadata
33
  filtering, with graceful semantic fallback when the company is ambiguous.
@@ -41,13 +47,14 @@ citations** — instead of making things up.
41
  ## 🏗️ Architecture
42
 
43
  ```
44
- INGESTION (once)
45
- fetch recent 10-Ks from SEC EDGAR ──► split into chunks
46
- ──► embed (local model) ──► store vectors + metadata in ChromaDB
47
 
48
  QUERY (per question)
49
- question ──► detect company ──► embed ──► search (filtered by company)
50
- ──► top-k chunks ──► LLM ──► grounded answer + citations
 
51
  ```
52
 
53
  | Layer | Tool |
@@ -57,7 +64,7 @@ QUERY (per question)
57
  | Vector store | ChromaDB (persisted locally) |
58
  | LLM | Llama 3.3 70B via Groq (free) |
59
  | UI | Streamlit |
60
- | Data | Recent SEC 10-K filings via `edgartools` (EDGAR) |
61
  | Evaluation | Capability gold set + FinanceBench, LLM-as-judge |
62
 
63
  **Corpus — 25 recognizable companies (FY2021–2024 10-Ks):** Apple, Microsoft,
@@ -140,8 +147,9 @@ committed index.
140
 
141
  ## ⚠️ Limitations
142
 
143
- - **Numeric/analytical** questions (ratios, margins) require computation over
144
- financial-statement tables — the main gap (see Evaluation).
 
145
  - The corpus is scoped to 25 companies' recent 10-Ks to stay laptop-friendly.
146
  - Not financial advice — a portfolio/educational project.
147
 
 
28
 
29
  - **Grounded answers with citations** — every response is backed by excerpts
30
  from real 10-K filings, shown in an expandable *Sources* panel.
31
+ - **Answers financial figures (hybrid RAG)** — plain text RAG can't read numbers
32
+ out of financial-statement tables. FinChat also extracts each filing's
33
+ **XBRL** structured financials (revenue, net income, assets, cash flow…) and
34
+ a hybrid retriever *guarantees* those facts are in context for numeric
35
+ questions — so *"What was Apple's FY2023 revenue?"* returns **$383.29 billion**,
36
+ not a shrug.
37
  - **Query routing ("knows where to look")** — FinChat detects which company a
38
  question is about and searches *only* that company's filings via metadata
39
  filtering, with graceful semantic fallback when the company is ambiguous.
 
47
  ## 🏗️ Architecture
48
 
49
  ```
50
+ INGESTION (once) — two tracks per filing
51
+ 10-K TEXT ──► split into chunks ───────────┐
52
+ XBRL FINANCIALS ──► "label: value" fact chunks ──┴─► embed ──► ChromaDB
53
 
54
  QUERY (per question)
55
+ question ──► detect company ──► HYBRID retrieve
56
+ (semantic text chunks + guaranteed XBRL statements for numeric Qs)
57
+ ──► LLM ──► grounded answer + citations
58
  ```
59
 
60
  | Layer | Tool |
 
64
  | Vector store | ChromaDB (persisted locally) |
65
  | LLM | Llama 3.3 70B via Groq (free) |
66
  | UI | Streamlit |
67
+ | Data | SEC 10-K text **+ XBRL financials** via `edgartools` |
68
  | Evaluation | Capability gold set + FinanceBench, LLM-as-judge |
69
 
70
  **Corpus — 25 recognizable companies (FY2021–2024 10-Ks):** Apple, Microsoft,
 
147
 
148
  ## ⚠️ Limitations
149
 
150
+ - **Direct financial figures** (revenue, net income, assets, cash flow) are now
151
+ answered from XBRL data. **Computed metrics** (ratios, margin trends,
152
+ EBITDA-less-capex) still need a calculation layer — the next enhancement.
153
  - The corpus is scoped to 25 companies' recent 10-Ks to stay laptop-friendly.
154
  - Not financial advice — a portfolio/educational project.
155
 
src/financials.py ADDED
@@ -0,0 +1,114 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Extract structured financial figures from a 10-K's XBRL data.
2
+
3
+ SEC filings carry machine-readable XBRL financials. We pull the income
4
+ statement, balance sheet, and cash-flow statement as clean "label: value"
5
+ facts and index them alongside the filing text, so FinChat can answer numeric
6
+ questions (revenue, net income, assets, ...) that plain text RAG cannot -- the
7
+ tables in the filing text collapse into unusable "number soup".
8
+ """
9
+ from __future__ import annotations
10
+
11
+ import re
12
+
13
+ import pandas as pd
14
+ from langchain_core.documents import Document
15
+
16
+ _STATEMENTS = [
17
+ ("income_statement", "Income Statement"),
18
+ ("balance_sheet", "Balance Sheet"),
19
+ ("cash_flow_statement", "Cash Flow Statement"),
20
+ ]
21
+ # Synonyms in the chunk header so numeric queries retrieve these chunks.
22
+ _KEYWORDS = ("revenue, sales, income, earnings, profit, margin, assets, "
23
+ "liabilities, equity, cash flow, expenses, EPS")
24
+
25
+
26
+ def _fmt(value, label: str = "") -> str | None:
27
+ try:
28
+ v = float(value)
29
+ except (TypeError, ValueError):
30
+ return None
31
+ if pd.isna(v):
32
+ return None
33
+ a = abs(v)
34
+ if "shares" in label.lower() and "per share" not in label.lower():
35
+ if a >= 1e9:
36
+ return f"{v / 1e9:,.2f} billion shares"
37
+ if a >= 1e6:
38
+ return f"{v / 1e6:,.1f} million shares"
39
+ return f"{v:,.0f} shares"
40
+ if a >= 1e9:
41
+ return f"${v / 1e9:,.2f} billion"
42
+ if a >= 1e6:
43
+ return f"${v / 1e6:,.1f} million"
44
+ if a >= 1000:
45
+ return f"${v:,.0f}"
46
+ return f"{v:,.2f}"
47
+
48
+
49
+ def _value_column(df: pd.DataFrame, fiscal_year: int):
50
+ """The statement column for the filing's primary fiscal year."""
51
+ date_cols = [c for c in df.columns if re.match(r"\d{4}-\d{2}-\d{2}", str(c))]
52
+ for c in date_cols:
53
+ if str(c).startswith(str(fiscal_year)):
54
+ return c
55
+ return date_cols[0] if date_cols else None
56
+
57
+
58
+ def _statement_lines(stmt, fiscal_year: int) -> list[str]:
59
+ try:
60
+ df = stmt.to_dataframe()
61
+ except Exception:
62
+ return []
63
+ col = _value_column(df, fiscal_year)
64
+ if col is None:
65
+ return []
66
+ lines = []
67
+ for _, row in df.iterrows():
68
+ # skip section headers, per-dimension breakdowns, sub-members
69
+ if row.get("abstract") or row.get("is_breakdown") or row.get("dimension"):
70
+ continue
71
+ label = str(row.get("label") or "").strip()
72
+ if not label:
73
+ continue
74
+ value = _fmt(row.get(col), label)
75
+ if value is None:
76
+ continue
77
+ lines.append(f"{label}: {value}")
78
+ return lines
79
+
80
+
81
+ def financial_documents(filing, ticker: str, company: str,
82
+ fiscal_year: int) -> list[Document]:
83
+ """One Document per financial statement, built from the filing's XBRL data."""
84
+ try:
85
+ tenk = filing.obj()
86
+ except Exception:
87
+ return []
88
+
89
+ docs: list[Document] = []
90
+ for attr, name in _STATEMENTS:
91
+ stmt = getattr(tenk, attr, None)
92
+ if stmt is None:
93
+ continue
94
+ lines = _statement_lines(stmt, fiscal_year)
95
+ if not lines:
96
+ continue
97
+ header = (
98
+ f"{company} ({ticker}) FY{fiscal_year} {name} "
99
+ f"(financial figures: {_KEYWORDS}; from SEC XBRL data):"
100
+ )
101
+ docs.append(
102
+ Document(
103
+ page_content=header + "\n" + "\n".join(lines),
104
+ metadata={
105
+ "ticker": ticker,
106
+ "company": company,
107
+ "year": str(fiscal_year),
108
+ "accession": filing.accession_no,
109
+ "source": f"{company} 10-K (FY{fiscal_year}) - {name} (XBRL)",
110
+ "type": "financials",
111
+ },
112
+ )
113
+ )
114
+ return docs
src/ingest.py CHANGED
@@ -24,6 +24,7 @@ from langchain_huggingface import HuggingFaceEmbeddings
24
  from langchain_chroma import Chroma
25
 
26
  from src import config
 
27
 
28
  # Windows consoles default to cp1252 and crash when print() emits Unicode
29
  # (arrows, em-dashes, curly quotes from filings). Force UTF-8 output.
@@ -33,8 +34,8 @@ except Exception:
33
  pass
34
 
35
 
36
- def fetch_filing(ticker: str, fiscal_year: int) -> tuple[str, str] | None:
37
- """Return (text, accession_no) for the 10-K whose fiscal period matches.
38
 
39
  Matches on the filing's period_of_report year, so an offset fiscal year
40
  (e.g. Amcor's June close) still resolves to the right filing.
@@ -42,12 +43,12 @@ def fetch_filing(ticker: str, fiscal_year: int) -> tuple[str, str] | None:
42
  for f in Company(ticker).get_filings(form="10-K"):
43
  period = getattr(f, "period_of_report", None)
44
  if period and str(period)[:4] == str(fiscal_year):
45
- return f.text(), f.accession_no
46
  return None
47
 
48
 
49
  def fetch_documents() -> list[Document]:
50
- """Fetch every target filing and split it into chunked LangChain Documents."""
51
  set_identity(config.EDGAR_IDENTITY)
52
  splitter = RecursiveCharacterTextSplitter(
53
  chunk_size=config.CHUNK_SIZE,
@@ -57,11 +58,13 @@ def fetch_documents() -> list[Document]:
57
  docs: list[Document] = []
58
  for ticker, name, year in config.TARGET_FILINGS:
59
  print(f"Fetching {ticker} FY{year} 10-K ...", end=" ", flush=True)
60
- result = fetch_filing(ticker, year)
61
- if result is None:
62
  print("NOT FOUND -- skipping")
63
  continue
64
- text, accession = result
 
 
65
  source = f"{name} 10-K (FY{year})"
66
  for chunk in splitter.split_text(text):
67
  docs.append(
@@ -71,12 +74,19 @@ def fetch_documents() -> list[Document]:
71
  "ticker": ticker,
72
  "company": name,
73
  "year": str(year),
74
- "accession": accession,
75
  "source": source,
 
76
  },
77
  )
78
  )
79
- print(f"{len(text):,} chars -> {len(docs)} chunks so far")
 
 
 
 
 
 
80
  time.sleep(0.5) # be polite to SEC's servers
81
  return docs
82
 
 
24
  from langchain_chroma import Chroma
25
 
26
  from src import config
27
+ from src.financials import financial_documents
28
 
29
  # Windows consoles default to cp1252 and crash when print() emits Unicode
30
  # (arrows, em-dashes, curly quotes from filings). Force UTF-8 output.
 
34
  pass
35
 
36
 
37
+ def get_filing(ticker: str, fiscal_year: int):
38
+ """Return the 10-K filing whose fiscal period matches fiscal_year, else None.
39
 
40
  Matches on the filing's period_of_report year, so an offset fiscal year
41
  (e.g. Amcor's June close) still resolves to the right filing.
 
43
  for f in Company(ticker).get_filings(form="10-K"):
44
  period = getattr(f, "period_of_report", None)
45
  if period and str(period)[:4] == str(fiscal_year):
46
+ return f
47
  return None
48
 
49
 
50
  def fetch_documents() -> list[Document]:
51
+ """Fetch every target filing -> text chunks + structured financial facts."""
52
  set_identity(config.EDGAR_IDENTITY)
53
  splitter = RecursiveCharacterTextSplitter(
54
  chunk_size=config.CHUNK_SIZE,
 
58
  docs: list[Document] = []
59
  for ticker, name, year in config.TARGET_FILINGS:
60
  print(f"Fetching {ticker} FY{year} 10-K ...", end=" ", flush=True)
61
+ filing = get_filing(ticker, year)
62
+ if filing is None:
63
  print("NOT FOUND -- skipping")
64
  continue
65
+
66
+ # 1) Filing TEXT -> chunks (for qualitative questions).
67
+ text = filing.text()
68
  source = f"{name} 10-K (FY{year})"
69
  for chunk in splitter.split_text(text):
70
  docs.append(
 
74
  "ticker": ticker,
75
  "company": name,
76
  "year": str(year),
77
+ "accession": filing.accession_no,
78
  "source": source,
79
+ "type": "text",
80
  },
81
  )
82
  )
83
+
84
+ # 2) XBRL FINANCIAL STATEMENTS -> structured facts (numeric questions).
85
+ fin_docs = financial_documents(filing, ticker, name, year)
86
+ docs.extend(fin_docs)
87
+
88
+ print(f"{len(text):,} chars + {len(fin_docs)} financial statements "
89
+ f"-> {len(docs)} docs so far")
90
  time.sleep(0.5) # be polite to SEC's servers
91
  return docs
92
 
src/rag.py CHANGED
@@ -146,13 +146,43 @@ def detect_ticker(question: str) -> str | None:
146
  return None
147
 
148
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
149
  def retrieve(question: str, ticker: str | None):
150
  store = get_vectorstore()
151
  search_kwargs: dict = {"k": config.TOP_K}
152
  if ticker:
153
  # Metadata filter = search ONLY that company's filings.
154
  search_kwargs["filter"] = {"ticker": ticker}
155
- return store.as_retriever(search_kwargs=search_kwargs).invoke(question)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
156
 
157
 
158
  def format_context(docs) -> str:
 
146
  return None
147
 
148
 
149
+ # Terms that signal a numeric/financial question -> pull in the XBRL statements.
150
+ _FINANCIAL_TERMS = (
151
+ "revenue", "sales", "income", "earnings", "profit", "margin", "ebitda",
152
+ "asset", "liabilit", "equity", "cash flow", "cash", "debt", "expense",
153
+ "eps", "per share", "how much", "dividend", "operating", "gross", "net ",
154
+ "balance sheet", "capital", "ratio",
155
+ )
156
+
157
+
158
+ def _is_financial_query(question: str) -> bool:
159
+ q = question.lower()
160
+ return any(term in q for term in _FINANCIAL_TERMS)
161
+
162
+
163
  def retrieve(question: str, ticker: str | None):
164
  store = get_vectorstore()
165
  search_kwargs: dict = {"k": config.TOP_K}
166
  if ticker:
167
  # Metadata filter = search ONLY that company's filings.
168
  search_kwargs["filter"] = {"ticker": ticker}
169
+ docs = store.as_retriever(search_kwargs=search_kwargs).invoke(question)
170
+
171
+ # Hybrid step: for a numeric/financial question about a known company,
172
+ # guarantee that company's structured XBRL statements are in context --
173
+ # they can otherwise be out-ranked by revenue *discussion* in the filing
174
+ # text (as happens for Apple).
175
+ if ticker and _is_financial_query(question):
176
+ fin = store.as_retriever(
177
+ search_kwargs={
178
+ "k": 2,
179
+ "filter": {"$and": [{"ticker": ticker}, {"type": "financials"}]},
180
+ }
181
+ ).invoke(question)
182
+ seen = {d.page_content[:80] for d in fin}
183
+ rest = [d for d in docs if d.page_content[:80] not in seen]
184
+ docs = (fin + rest)[: config.TOP_K]
185
+ return docs
186
 
187
 
188
  def format_context(docs) -> str:
vectorstore/{615fd89b-a274-475a-8189-64a3c37ef9f7 → 106fd921-0890-496a-95dd-cba23b02f735}/data_level0.bin RENAMED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:fb5e357d3efb20827f7972d9d72b0384f9301878ddd40f2d1479c5c82388bfa8
3
- size 35541256
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:8c928a39e9351d4c9febd587413bbafab14c4a987e8ebee394f08804697a4cab
3
+ size 35666956
vectorstore/{615fd89b-a274-475a-8189-64a3c37ef9f7 → 106fd921-0890-496a-95dd-cba23b02f735}/header.bin RENAMED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:0093aec8945d3efc9942dcf58966ec769316b3f1cb51497648bb773d43311c6e
3
  size 100
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a5f9fd9f1648dd4e3fdfec3fa226de568f9c5af5f9bea7d5a42b7b71dddab8b7
3
  size 100
vectorstore/{615fd89b-a274-475a-8189-64a3c37ef9f7 → 106fd921-0890-496a-95dd-cba23b02f735}/index_metadata.pickle RENAMED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:a886d4dae7bf3edd8aca30a00285319ef1bac904a694169ff526fe65ccacab67
3
- size 1951164
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:3bcf185e708e3adb78620b53ea97246cb1d26fec4e928749e570bd2bab6d938c
3
+ size 1958064
vectorstore/{615fd89b-a274-475a-8189-64a3c37ef9f7 → 106fd921-0890-496a-95dd-cba23b02f735}/length.bin RENAMED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:19f08c1fcb54ea3c693532da8f45031a8827042407de975fdd24bdf9424661dd
3
- size 84824
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f30b58007de341e2f2ba7cb41cd5c1865f6d69e5af157a1fae4e6bccb739c94e
3
+ size 85124
vectorstore/{615fd89b-a274-475a-8189-64a3c37ef9f7 → 106fd921-0890-496a-95dd-cba23b02f735}/link_lists.bin RENAMED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:1379a6c040ce26666a0cb509ceceacdcac173bc54f8199e661a9af5d776277f0
3
- size 182336
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b7a26b0bce7bcdf3f6e9b091cc73b09960b452087339c31b059a9a0cbd3eebf3
3
+ size 181956
vectorstore/chroma.sqlite3 CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:5aaad139c84897df5458309236ca3fa8200a68df2fe85b5d550d87921d97cd92
3
- size 122077184
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:4e7d02066f6535862c4db2e8ef3ce5ac23ecb2dee1a1d8e16468a7f4603b489a
3
+ size 124530688