Spaces:
Running on Zero
Download README.md from Tim-Pinecone/SEC-FTS-Hybrid-Comparison: direct link, hf CLI and curl.
- Browser
- Download file 8.78 kB
-
https://huggingface.co/spaces/Tim-Pinecone/SEC-FTS-Hybrid-Comparison/resolve/main/README.md
- Command line
-
hf download hf://spaces/Tim-Pinecone/SEC-FTS-Hybrid-Comparison/README.md
-
curl -L -o README.md https://huggingface.co/spaces/Tim-Pinecone/SEC-FTS-Hybrid-Comparison/resolve/main/README.md
A newer version of the Gradio SDK is available: 6.29.1
title: SEC FTS Hybrid Comparison
emoji: π
colorFrom: blue
colorTo: indigo
sdk: gradio
sdk_version: 6.29.0
python_version: '3.12'
app_file: app.py
pinned: false
short_description: Pinecone full-text search vs dense vectors on 10-Ks
SEC Document Search β Pinecone FTS + Semantic
A demo application showing how to build a hybrid search system over SEC 10-K filings using Pinecone's full-text search combined with dense vector embeddings.
The app supports three search modes:
| Mode | How it works |
|---|---|
| Full-text | BM25 keyword search over the document text |
| Query builder (full-text) | Construct complex full-text queries β Lucene boolean / phrase / proximity / boost / prefix scoring plus text-match and metadata filters β starting from 10 worked examples |
| Semantic | Dense vector similarity via OpenAI embeddings |
| Compare: dense vs full-text | Runs a dense search and a full-text search side by side and contrasts them β overlap, rank shifts, keyword coverage, latency, and an RRF-fused list |
| Hybrid | Dense vector ranking with a must-contain keyword filter β results are semantically ranked and guaranteed to contain the specified terms |
Data
example_data/ contains chunked 10-K filings for six companies across six years:
| Ticker | Company | Years |
|---|---|---|
| AAPL | Apple | 2019β2024 |
| AMZN | Amazon | 2019β2024 |
| F | Ford | 2019β2024 |
| GM | General Motors | 2019β2024 |
| MSFT | Microsoft | 2019β2024 |
| ORCL | Oracle | 2019β2024 |
~24,000 document chunks total. Each chunk has an _id, text, ticker, filing_type, year, and chunk_index.
Prerequisites
- Python 3.12+
uv- A Pinecone API key (free tier works)
- An OpenAI API key
Setup
git clone https://github.com/tim-pinecone/sec-dense-fts
cd sec-dense-fts
cp .env.example .env
# Add your PINECONE_API_KEY and OPENAI_API_KEY to .env
uv sync
Ingest
Creates the Pinecone index, embeds all chunks with text-embedding-3-small, and upserts them in batches. Safe to re-run β skips index creation if it already exists.
uv run main.py
Ingestion takes a few minutes (OpenAI embedding calls are the bottleneck). The script polls until all documents are searchable before exiting.
Run the app
Two UIs share the same search logic (search_core.py, fts_queries.py):
| UI | Command | URL | Notes |
|---|---|---|---|
| Gradio | uv run python app.py |
http://localhost:7860 | What the Hugging Face Space runs |
| Streamlit | uv run streamlit run app_streamlit.py |
http://localhost:8501 | Local development UI |
Both have the same five tabs/modes. Use the sidebar to filter by ticker and year, pick a search mode, and enter your query β or pick one of the prepared Example queries at the top of each mode (three per mode; ten in the query builder). Examples live in fts_queries.py (MODE_EXAMPLES, EXAMPLES).
Full-text query builder
The Query builder mode exposes the full Pinecone FTS query surface. Queries are built in three parts, and the exact documents.search(...) request is shown as JSON and Python before running:
- Scoring β either BM25 keywords (
text) or Lucene (query_string). Lucene can be written raw or assembled row-by-row with the clause builder (MUST+/ MUST NOT-/ SHOULD, term / phrase / phrase prefix, slop~N, boost^N). - Text-match filters β
$match_phrase,$match_all,$match_any, each optionally negated with$not, combined with$andor$or. - Metadata filters β ticker
$in/$nin, year range, chunk-index range.
Pick an example from the dropdown to load it into the builder, then tweak it:
| Example | Demonstrates |
|---|---|
| Supply-chain shortages, excluding COVID | text:(("supply chain" OR semiconductor) AND shortage) NOT text:(covid) |
| Cyber incidents | text:(+cybersecurity ransomware^3 breach -insurance) |
| Rising interest rates | Proximity: text:("interest rates increase"~5) |
| China trade & tariffs | Term boost: text:(tariffs^3 trade china) |
| AI mentions | Phrase prefix "artificial intel"* + required term + year range |
| EV batteries at Ford & GM | BM25 + $match_phrase filter + ticker $in + year range |
| Cloud growth, no pandemic talk | $not + $match_any exclusion, ticker $nin |
| Regulators: EC or DOJ | $or across two $match_phrase filters |
| Inflation & input costs | $match_all + year range |
| Buybacks vs. dividends | Required OR-groups +(a OR b) +(c OR d) |
Things the server enforces (surfaced in the UI):
- One scoring type per request β
textorquery_string, never mixed. query_stringclauses may not setfields; qualify terms inline (text:(...)).- A Lucene query of only exclusions (
text:(-covid)) is rejected. - Proximity, boost and phrase prefix are scoring-only β they can't be used in
filter. - Phrase-prefix matches all receive the same constant score.
The query compilation logic and examples live in fts_queries.py.
Comparing dense vs full-text
Compare mode runs both searches in parallel for the same question:
- Dense side β the query is embedded with
text-embedding-3-smalland ranked by cosine similarity. - Full-text side β either the same query as BM25 keywords, or whatever complex query is currently set up in the Query builder (Lucene, text-match filters, metadata filters). When using the builder, you can optionally apply its filters to the dense side too, so only the ranking signal differs.
Sidebar ticker/year filters apply to both sides. The results show:
| View | What it tells you |
|---|---|
| Overlap / Jaccard, only-dense, only-full-text | How much the two retrieval methods agree in the top-k |
| Latency | Dense (embedding + search) vs full-text search time |
| Keyword coverage | Share of the query's keywords present in each result β dense results with low coverage are paraphrase / concept matches that BM25 can't find |
| Side by side | Both ranked lists, with badges showing each result's rank in the other list, and keyword highlighting on both |
| Rank comparison | One table of every retrieved chunk with its dense rank, full-text rank and Ξ |
| Fused (RRF) | Client-side reciprocal rank fusion (k = 60) of the two lists β a preview of a two-query hybrid |
Deploying to Hugging Face Spaces
The Space runs the Gradio app (app.py) on free ZeroGPU hardware β no Docker needed. The YAML block at the top of this README is the Space config, and requirements.txt holds the Space's Python dependencies (Gradio itself comes from sdk_version). ZeroGPU requires at least one @spaces.GPU function; app.py defines a no-op one since all compute happens in Pinecone and OpenAI.
- Add a Hugging Face write token to
.envasHF_TOKEN. - Preview what will be uploaded:
uv run python deploy_space.py <owner>/<space-name> --dry-run - Deploy. The first run creates the Space;
--set-secretscopiesPINECONE_API_KEYandOPENAI_API_KEYfrom.envinto the Space's secrets (only needed once, or when keys change):uv run python deploy_space.py <owner>/<space-name> --set-secrets [--private]
deploy_space.py uploads an explicit allowlist (README.md, requirements.txt, app.py, search_core.py, fts_queries.py), so .env, the example data, and the Streamlit app never leave your machine. Re-run it without --set-secrets to push code changes.
Index schema
text β full-text search field (BM25, English, stemming enabled)
embedding β dense vector, 1536 dims, cosine similarity
ticker β filterable metadata (string)
filing_type β filterable metadata (string)
year β filterable metadata (integer)
chunk_index β filterable metadata (integer)
The FTS and vector fields are declared in the schema at index creation. Metadata fields (ticker, filing_type, year, chunk_index) are automatically indexed β they do not need to be declared.
How hybrid search works
The hybrid mode uses a single Pinecone query:
score_byβ dense vector cosine similarity (semantic ranking)filterβ$match_allon the text field (hard lexical requirement)
This means results are ordered by semantic relevance, but only chunks that contain all the specified keywords are returned. It's useful for queries like "what does MSFT say about Azure capital expenditure" β the semantic query captures the intent, and the text filter ensures the specific terms are present.