Spaces:
Sleeping
Download README.md from jashp2323/visual-search: direct link, hf CLI and curl.
- Browser
- Download file 9.42 kB
-
https://huggingface.co/spaces/jashp2323/visual-search/resolve/main/README.md
- Command line
-
hf download hf://spaces/jashp2323/visual-search/README.md
-
curl -L -o README.md https://huggingface.co/spaces/jashp2323/visual-search/resolve/main/README.md
title: AURA Neural Search
emoji: 🔍
colorFrom: indigo
colorTo: purple
sdk: docker
app_port: 7860
pinned: false
AURA // Neural Visual Search Engine
A high-performance, fully local multi-modal visual search engine matching Pinterest's visual search and Google Lens. Built with FastAPI, CLIP (Contrastive Language-Image Pre-training), and NumPy.
The search engine features a glassmorphism frontend and works with a high-resolution catalog of 100 fashion products streamed directly from Kaggle's fashion dataset mirror on Hugging Face.
Setup & Running Guide
1. Prerequisites
- Python 3.9 or higher
- Recommended: uv (fast Python package installer and manager)
2. Setup Virtual Environment & Install Dependencies
Using standard pip:
# Create virtual environment
python -m venv .venv
# Activate virtual environment
.venv\Scripts\activate
# Install requirements
pip install -r requirements.txt
Using uv (Recommended / Ultra Fast):
# Create virtual environment
uv venv
# Install requirements
uv pip install -r requirements.txt
3. Initialize the Catalog & Precompute Embeddings
Run the catalog generator script. This will download 100 high-resolution fashion products (apparel, footwear, accessories, bags) and precompute their CLIP embeddings into catalog.json and embeddings.npy:
.venv\Scripts\python.exe generate_catalog.py
4. Run the Web Server
Launch the FastAPI development server using Uvicorn:
.venv\Scripts\python.exe -m uvicorn app:app --host 127.0.0.1 --port 8000 --reload
5. Open the Web Application
Open your browser and navigate to: http://127.0.0.1:8000/
Production Deployment (Render & Vercel)
This application is configured to run in a split deployment configuration:
- Backend: Deployed on Render (as a Docker Web Service).
- Frontend: Deployed on Vercel (hosting the static client assets).
1. Backend Deployment on Render
Render will automatically build and run the backend using the provided Dockerfile (optimized for CPU to minimize build size and time).
- Go to your Render Dashboard and click New > Web Service.
- Connect your GitHub repository (
jashparmaraiml-20/Visual-Search). - Select Docker as the Runtime (Render will automatically detect the
Dockerfileat the root). - Set the Instance Type to Free (or higher).
- Render will automatically inject the
$PORTenvironment variable. The Dockerfile's launch command is configured to bind to0.0.0.0:$PORTautomatically. - Once deployed, note down your Render Web Service URL (e.g.,
https://visual-search-backend.onrender.com).
2. Frontend Configuration & Vercel Deployment
Before deploying the frontend, update static/app.js with your backend endpoint.
- Open
static/app.jsand set theAPI_BASEvariable at the top of the file to your Render backend URL:const API_BASE = "https://your-backend.onrender.com"; // Replace with your Render URL - Commit and push this change to your repository.
- Go to Vercel and import your project from GitHub.
- Set the Root Directory of the project to
static(so Vercel only serves the static client assets:index.html,app.js,style.css, andimages/). - Click Deploy. Vercel will host the frontend on its fast edge CDN.
System Design & Core Concepts
1. Vector Embeddings with CLIP
At the heart of the system is the CLIP model (openai/clip-vit-base-patch32), which aligns images and text in the same semantic space of 512 dimensions.
- Text Embedding: A text query (e.g., "red sneakers") is tokenized and encoded by the text transformer into a 512-dimensional vector.
- Image Embedding: An image is resized, normalized, and encoded by the vision transformer into a 512-dimensional vector.
- Cosine Similarity: The cosine similarity between any two vectors $u$ and $v$ is calculated as: $$\text{Similarity}(u, v) = \frac{u \cdot v}{|u| |v|}$$ Since we normalize vectors to unit length ($|u| = 1$), the cosine similarity is simplified to the dot product: $$\text{Similarity}(u, v) = u \cdot v$$ This allows lightning-fast search using simple matrix multiplication.
2. High-Level Architecture
The system consists of three main components:
- Premium Frontend UI: Single Page Application implementing a glassmorphism theme, drag-and-drop image upload, instant search feedback, and instant visual similarity queries ("Find Similar" hover feature).
- FastAPI Web Server: Serves API routes, runs the local CLIP inference engine, manages the product catalog, and serves static files (product images, HTML, CSS, JS).
- Local Vector Store: Memory-resident vector database using NumPy matrix calculations.
graph TD
A[Client Browser UI] -->|Text Query / Image Upload| B[FastAPI Backend]
B -->|Preprocess Input| C[CLIP Model Inference]
C -->|Generate Embedding Vector| D[Vector Matcher]
D -->|Cosine Similarity Dot Product| E[(Local Vector Store / Catalog)]
E -->|Sorted Match Results| B
B -->|JSON Response| A
Detailed Data Flows & Userflows
Userflow 1: Text-to-Image Search (e.g., searching "red sneakers")
sequenceDiagram
autonumber
actor User
participant UI as Frontend (JS/CSS)
participant API as FastAPI Server
participant CLIP as CLIP Inference Engine
participant DB as Local Vector Store (NumPy)
User->>UI: Types "red sneakers" and clicks Search/presses Enter
UI->>API: POST /api/search (Form Data: q="red sneakers")
API->>CLIP: Generate text embedding (shape: [1, 512])
CLIP-->>API: Return embedding vector
API->>DB: Query similarity (dot product against catalog embeddings matrix [N, 512])
DB-->>API: Sorted index and similarity scores
API-->>UI: JSON response: [{product_id, title, score, image_url}, ...]
UI->>User: Render matches with smooth fade-in animations
Userflow 2: Drag-and-Drop Image Visual Search
sequenceDiagram
autonumber
actor User
participant UI as Frontend (JS/CSS)
participant API as FastAPI Server
participant CLIP as CLIP Inference Engine
participant DB as Local Vector Store (NumPy)
User->>UI: Drags & drops or uploads an image (e.g., of a leather jacket)
UI->>API: POST /api/search (Multipart Form Data: file)
API->>CLIP: Generate image embedding (shape: [1, 512])
CLIP-->>API: Return embedding vector
API->>DB: Query similarity (dot product against catalog embeddings matrix [N, 512])
DB-->>API: Sorted index and similarity scores
API-->>UI: JSON response: [{product_id, title, score, image_url}, ...]
UI->>User: Update results showing visually matching items
Userflow 3: "Find Similar" (Pinterest style)
sequenceDiagram
autonumber
actor User
participant UI as Frontend (JS/CSS)
participant API as FastAPI Server
participant DB as Local Vector Store (NumPy)
User->>UI: Hovers product card, clicks "Find Similar" button
UI->>API: POST /api/search/similar (JSON: {product_id: "prod-4"})
API->>DB: Retrieve cached embedding for "prod-4" (no CLIP run needed!)
API->>DB: Dot product target embedding against all other catalog vectors
DB-->>API: Sorted similarity rankings (excluding self)
API-->>UI: JSON response: visually similar product catalog list
UI->>User: Re-render results grid with similar products
Component Details
1. Vector Store (vector_store.py)
Encapsulates vector insertion and query logic.
- Data Structure: Standard Python
dictmapping product IDs to product metadata, and a NumPy array storing normalized vector embeddings. - Search Function: Calculates cosine similarity using raw matrix-vector multiplication (
np.dot(embeddings, query_vector)). - Scaling / 1-Line Swap:
# Simple switch in vector_store.py # self.provider = InMemoryVectorStore() # self.provider = QdrantVectorStore() # <-- SWAP HERE to scale to millions of items
2. CLIP Model Wrapper (models.py)
Handles downloading, caching, and inference of the CLIP model using transformers and torch.
- Model:
openai/clip-vit-base-patch32(balanced size, fast CPU execution, high-accuracy multi-modal embeddings). - Device: Auto-selects CUDA GPU if available; otherwise defaults to CPU.
- Methods:
get_text_embedding(text: str) -> np.ndarrayget_image_embedding(image: PIL.Image) -> np.ndarray
3. Catalog Generator (generate_catalog.py)
Initializes the system by downloading fashion product images and computing their embeddings.
- Streams high-res (384x512) product data from Hugging Face.
- Precomputes product visual embeddings and saves them to enabling instant server startup.
4. FastAPI Server (app.py)
Defines the HTTP interface:
GET /: Serves UI.GET /api/products: Fetches all catalog products.POST /api/search: Multi-modal search. Handles either text query or uploaded image file.POST /api/search/similar: Look up cached embedding for a given product ID and return matching items.POST /api/index: Adds a new product to the catalog, downloads image, computes embedding, and saves to vector store.