--- title: AURA Neural Search emoji: 🔍 colorFrom: indigo colorTo: purple sdk: docker app_port: 7860 pinned: false --- # AURA // Neural Visual Search Engine A high-performance, fully local multi-modal visual search engine matching Pinterest's visual search and Google Lens. Built with **FastAPI**, **CLIP** (Contrastive Language-Image Pre-training), and **NumPy**. The search engine features a glassmorphism frontend and works with a high-resolution catalog of 100 fashion products streamed directly from Kaggle's fashion dataset mirror on Hugging Face. --- ## Setup & Running Guide ### 1. Prerequisites * Python 3.9 or higher * Recommended: [uv](https://github.com/astral-sh/uv) (fast Python package installer and manager) ### 2. Setup Virtual Environment & Install Dependencies #### Using standard `pip`: ```powershell # Create virtual environment python -m venv .venv # Activate virtual environment .venv\Scripts\activate # Install requirements pip install -r requirements.txt ``` #### Using `uv` (Recommended / Ultra Fast): ```powershell # Create virtual environment uv venv # Install requirements uv pip install -r requirements.txt ``` ### 3. Initialize the Catalog & Precompute Embeddings Run the catalog generator script. This will download 100 high-resolution fashion products (apparel, footwear, accessories, bags) and precompute their CLIP embeddings into `catalog.json` and `embeddings.npy`: ```powershell .venv\Scripts\python.exe generate_catalog.py ``` ### 4. Run the Web Server Launch the FastAPI development server using Uvicorn: ```powershell .venv\Scripts\python.exe -m uvicorn app:app --host 127.0.0.1 --port 8000 --reload ``` ### 5. Open the Web Application Open your browser and navigate to: **[http://127.0.0.1:8000/](http://127.0.0.1:8000/)** --- ## Production Deployment (Render & Vercel) This application is configured to run in a split deployment configuration: - **Backend**: Deployed on **Render** (as a Docker Web Service). - **Frontend**: Deployed on **Vercel** (hosting the static client assets). ### 1. Backend Deployment on Render Render will automatically build and run the backend using the provided `Dockerfile` (optimized for CPU to minimize build size and time). 1. Go to your **Render Dashboard** and click **New > Web Service**. 2. Connect your GitHub repository (`jashparmaraiml-20/Visual-Search`). 3. Select **Docker** as the Runtime (Render will automatically detect the `Dockerfile` at the root). 4. Set the Instance Type to **Free** (or higher). 5. Render will automatically inject the `$PORT` environment variable. The Dockerfile's launch command is configured to bind to `0.0.0.0:$PORT` automatically. 6. Once deployed, note down your Render Web Service URL (e.g., `https://visual-search-backend.onrender.com`). ### 2. Frontend Configuration & Vercel Deployment Before deploying the frontend, update `static/app.js` with your backend endpoint. 1. Open `static/app.js` and set the `API_BASE` variable at the top of the file to your Render backend URL: ```javascript const API_BASE = "https://your-backend.onrender.com"; // Replace with your Render URL ``` 2. Commit and push this change to your repository. 3. Go to **Vercel** and import your project from GitHub. 4. Set the **Root Directory** of the project to `static` (so Vercel only serves the static client assets: `index.html`, `app.js`, `style.css`, and `images/`). 5. Click **Deploy**. Vercel will host the frontend on its fast edge CDN. --- ## System Design & Core Concepts ### 1. Vector Embeddings with CLIP At the heart of the system is the CLIP model (`openai/clip-vit-base-patch32`), which aligns images and text in the same semantic space of 512 dimensions. * **Text Embedding**: A text query (e.g., "red sneakers") is tokenized and encoded by the text transformer into a 512-dimensional vector. * **Image Embedding**: An image is resized, normalized, and encoded by the vision transformer into a 512-dimensional vector. * **Cosine Similarity**: The cosine similarity between any two vectors $u$ and $v$ is calculated as: $$\text{Similarity}(u, v) = \frac{u \cdot v}{\|u\| \|v\|}$$ Since we normalize vectors to unit length ($\|u\| = 1$), the cosine similarity is simplified to the dot product: $$\text{Similarity}(u, v) = u \cdot v$$ This allows lightning-fast search using simple matrix multiplication. ### 2. High-Level Architecture The system consists of three main components: 1. **Premium Frontend UI**: Single Page Application implementing a glassmorphism theme, drag-and-drop image upload, instant search feedback, and instant visual similarity queries ("Find Similar" hover feature). 2. **FastAPI Web Server**: Serves API routes, runs the local CLIP inference engine, manages the product catalog, and serves static files (product images, HTML, CSS, JS). 3. **Local Vector Store**: Memory-resident vector database using NumPy matrix calculations. ```mermaid graph TD A[Client Browser UI] -->|Text Query / Image Upload| B[FastAPI Backend] B -->|Preprocess Input| C[CLIP Model Inference] C -->|Generate Embedding Vector| D[Vector Matcher] D -->|Cosine Similarity Dot Product| E[(Local Vector Store / Catalog)] E -->|Sorted Match Results| B B -->|JSON Response| A ``` --- ## Detailed Data Flows & Userflows ### Userflow 1: Text-to-Image Search (e.g., searching "red sneakers") ```mermaid sequenceDiagram autonumber actor User participant UI as Frontend (JS/CSS) participant API as FastAPI Server participant CLIP as CLIP Inference Engine participant DB as Local Vector Store (NumPy) User->>UI: Types "red sneakers" and clicks Search/presses Enter UI->>API: POST /api/search (Form Data: q="red sneakers") API->>CLIP: Generate text embedding (shape: [1, 512]) CLIP-->>API: Return embedding vector API->>DB: Query similarity (dot product against catalog embeddings matrix [N, 512]) DB-->>API: Sorted index and similarity scores API-->>UI: JSON response: [{product_id, title, score, image_url}, ...] UI->>User: Render matches with smooth fade-in animations ``` ### Userflow 2: Drag-and-Drop Image Visual Search ```mermaid sequenceDiagram autonumber actor User participant UI as Frontend (JS/CSS) participant API as FastAPI Server participant CLIP as CLIP Inference Engine participant DB as Local Vector Store (NumPy) User->>UI: Drags & drops or uploads an image (e.g., of a leather jacket) UI->>API: POST /api/search (Multipart Form Data: file) API->>CLIP: Generate image embedding (shape: [1, 512]) CLIP-->>API: Return embedding vector API->>DB: Query similarity (dot product against catalog embeddings matrix [N, 512]) DB-->>API: Sorted index and similarity scores API-->>UI: JSON response: [{product_id, title, score, image_url}, ...] UI->>User: Update results showing visually matching items ``` ### Userflow 3: "Find Similar" (Pinterest style) ```mermaid sequenceDiagram autonumber actor User participant UI as Frontend (JS/CSS) participant API as FastAPI Server participant DB as Local Vector Store (NumPy) User->>UI: Hovers product card, clicks "Find Similar" button UI->>API: POST /api/search/similar (JSON: {product_id: "prod-4"}) API->>DB: Retrieve cached embedding for "prod-4" (no CLIP run needed!) API->>DB: Dot product target embedding against all other catalog vectors DB-->>API: Sorted similarity rankings (excluding self) API-->>UI: JSON response: visually similar product catalog list UI->>User: Re-render results grid with similar products ``` --- ## Component Details ### 1. Vector Store (`vector_store.py`) Encapsulates vector insertion and query logic. * **Data Structure**: Standard Python `dict` mapping product IDs to product metadata, and a NumPy array storing normalized vector embeddings. * **Search Function**: Calculates cosine similarity using raw matrix-vector multiplication (`np.dot(embeddings, query_vector)`). * **Scaling / 1-Line Swap**: ```python # Simple switch in vector_store.py # self.provider = InMemoryVectorStore() # self.provider = QdrantVectorStore() # <-- SWAP HERE to scale to millions of items ``` ### 2. CLIP Model Wrapper (`models.py`) Handles downloading, caching, and inference of the CLIP model using `transformers` and `torch`. * **Model**: `openai/clip-vit-base-patch32` (balanced size, fast CPU execution, high-accuracy multi-modal embeddings). * **Device**: Auto-selects CUDA GPU if available; otherwise defaults to CPU. * **Methods**: * `get_text_embedding(text: str) -> np.ndarray` * `get_image_embedding(image: PIL.Image) -> np.ndarray` ### 3. Catalog Generator (`generate_catalog.py`) Initializes the system by downloading fashion product images and computing their embeddings. * Streams high-res (384x512) product data from Hugging Face. * Precomputes product visual embeddings and saves them to enabling instant server startup. ### 4. FastAPI Server (`app.py`) Defines the HTTP interface: * `GET /`: Serves UI. * `GET /api/products`: Fetches all catalog products. * `POST /api/search`: Multi-modal search. Handles either text query or uploaded image file. * `POST /api/search/similar`: Look up cached embedding for a given product ID and return matching items. * `POST /api/index`: Adds a new product to the catalog, downloads image, computes embedding, and saves to vector store.