# CLIP-based Local Visual Search Engine A high-performance, fully local visual search engine built with FastAPI and CLIP (Contrastive Language-Image Pre-training) matching Pinterest's visual search and Google Lens. --- ## System Design & Core Concepts ### 1. Vector Embeddings with CLIP At the heart of the system is the CLIP model (`openai/clip-vit-base-patch32`), which aligns images and text in the same semantic space of 512 dimensions. * **Text Embedding**: A text query (e.g., "red sneakers") is tokenized and encoded by the text transformer into a 512-dimensional vector. * **Image Embedding**: An image is resized, normalized, and encoded by the vision transformer into a 512-dimensional vector. * **Cosine Similarity**: The cosine similarity between any two vectors $u$ and $v$ is calculated as: $$\text{Similarity}(u, v) = \frac{u \cdot v}{\|u\| \|v\|}$$ Since we normalize vectors to unit length ($\|u\| = 1$), the cosine similarity is simplified to the dot product: $$\text{Similarity}(u, v) = u \cdot v$$ This allows lightning-fast search using simple matrix multiplication. ### 2. High-Level Architecture The system consists of three main components: 1. **Premium Frontend UI**: Single Page Application implementing a glassmorphism theme, dynamic CSS, drag-and-drop image upload, instant search feedback, and instant visual similarity queries ("Find Similar" hover feature). 2. **FastAPI Web Server**: Serves API routes, runs the local CLIP inference engine, manages the product catalog, and serves static files (product images, HTML, CSS, JS). 3. **Local Vector Store**: Memory-resident vector database using NumPy matrix calculations. Includes a configuration option to swap out this local engine for an external vector database (like Qdrant or Milvus) with just one line. ```mermaid graph TD A[Client Browser UI] -->|Text Query / Image Upload| B[FastAPI Backend] B -->|Preprocess Input| C[CLIP Model Inference] C -->|Generate Embedding Vector| D[Vector Matcher] D -->|Cosine Similarity Dot Product| E[(Local Vector Store / Catalog)] E -->|Sorted Match Results| B B -->|JSON Response with similarity scores| A ``` --- ## Detailed Data Flows & Userflows ### Userflow 1: Text-to-Image Search (e.g., searching "red sneakers") ```mermaid sequenceDiagram autonumber actor User participant UI as Frontend (JS/CSS) participant API as FastAPI Server participant CLIP as CLIP Inference Engine participant DB as Local Vector Store (NumPy) User->>UI: Types "red sneakers" and clicks Search/presses Enter UI->>API: POST /api/search (JSON: {query: "red sneakers"}) API->>CLIP: Generate text embedding (shape: [1, 512]) CLIP-->>API: Return embedding vector API->>DB: Query similarity (dot product against catalog embeddings matrix [N, 512]) DB-->>API: Sorted index and similarity scores (e.g. 0.82, 0.45...) API-->>UI: JSON response: [{product_id, title, score, image_url}, ...] UI->>User: Render matches with smooth fade-in animations and highlight score ``` ### Userflow 2: Drag-and-Drop Image Visual Search ```mermaid sequenceDiagram autonumber actor User participant UI as Frontend (JS/CSS) participant API as FastAPI Server participant CLIP as CLIP Inference Engine participant DB as Local Vector Store (NumPy) User->>UI: Drags & drops or uploads an image (e.g., of a leather jacket) UI->>API: POST /api/search (Multipart Form Data: file) API->>CLIP: Generate image embedding (shape: [1, 512]) CLIP-->>API: Return embedding vector API->>DB: Query similarity (dot product against catalog embeddings matrix [N, 512]) DB-->>API: Sorted index and similarity scores API-->>UI: JSON response: [{product_id, title, score, image_url}, ...] UI->>User: Update results showing visually matching items ``` ### Userflow 3: "Find Similar" (Pinterest style) ```mermaid sequenceDiagram autonumber actor User participant UI as Frontend (JS/CSS) participant API as FastAPI Server participant DB as Local Vector Store (NumPy) User->>UI: Hovers product card, clicks "Find Similar" button UI->>API: POST /api/search/similar (JSON: {product_id: "prod-4"}) API->>DB: Retrieve cached embedding for "prod-4" (no CLIP run needed!) API->>DB: Dot product target embedding against all other catalog vectors DB-->>API: Sorted similarity rankings (excluding self) API-->>UI: JSON response: visually similar product catalog list UI->>User: Re-render results grid with similar products ``` --- ## Component Details ### 1. Vector Store (`vector_store.py`) Encapsulates vector insertion and query logic. * **Data Structure**: Standard Python `dict` mapping product IDs to product metadata, and a NumPy array storing normalized vector embeddings. * **Search Function**: Calculates cosine similarity using raw matrix-vector multiplication (`np.dot(embeddings, query_vector)`). * **Scaling / 1-Line Swap**: ```python # Simple switch in vector_store.py # self.provider = InMemoryVectorStore() # self.provider = QdrantVectorStore() # <-- SWAP HERE to scale to millions of items ``` ### 2. CLIP Model Wrapper (`models.py`) Handles downloading, caching, and inference of the CLIP model using `transformers` and `torch`. * **Model**: `openai/clip-vit-base-patch32` (balanced size, fast CPU execution, high-accuracy multi-modal embeddings). * **Device**: Auto-selects CUDA GPU if available; otherwise defaults to CPU. * **Methods**: * `embed_text(text: str) -> np.ndarray` * `embed_image(image: PIL.Image) -> np.ndarray` ### 3. Catalog Generator (`generate_catalog.py`) Initializes the system by creating product images and computing their embeddings. * Uses local image files or generates a set of 10 high-quality aesthetic product images. * Precomputes product visual embeddings and saves them to a file (`catalog_with_embeddings.pkl` or `.json` + `.npy`) to enable near-instant server startup. ### 4. FastAPI Server (`app.py`) Defines the HTTP interface: * `GET /`: Serves UI. * `GET /api/products`: Fetches all catalog products. * `POST /api/search`: Multi-modal search. Handles either text query in JSON body OR uploaded image file. * `POST /api/search/similar`: Look up cached embedding for a given product ID and return matching items. * `POST /api/index`: Adds a new product to the catalog, downloads image, computes embedding, and saves to vector store. --- ## UI/UX & Styling Details To ensure a **premium experience that WOWs the user**: * **Theme**: Modern dark mode with glassmorphism panels (`backdrop-filter: blur(16px)`). * **Palette**: Curated deep colors using HSL (Hue, Saturation, Lightness) for unified tones: * Primary background: HSL(224, 25%, 8%) - Midnight navy * Cards/Panels: HSL(224, 20%, 14%, 0.7) - Semitransparent slate * Accent/Active: HSL(263, 70%, 50%) - Vibrant violet * Text Primary: HSL(0, 0%, 98%) - Off-white * **Transitions**: Smooth animations for hovering card expansions, searching states (pulsing glowing search rings), and layout sorting transitions. * **Layout**: CSS Grid with auto-fill columns. Responsive behavior from mobile to ultra-wide displays. * **Interactive visual crop / focus**: Product card shows similarity score (e.g., "94% Visual Match") in a styled badge. "Find Similar" action on card lets user drill down on an item. --- ## Proposed Directory & File Map * `[NEW] app.py` (FastAPI Server) * `[NEW] models.py` (CLIP Inference Wrapper) * `[NEW] vector_store.py` (Local NumPy Vector Store) * `[NEW] generate_catalog.py` (Helper script to generate mock images and build initial embed index) * `[NEW] requirements.txt` (Dependencies) * `[NEW] static/` * `[NEW] index.html` (Client UI) * `[NEW] style.css` (Glassmorphism layout design system) * `[NEW] app.js` (Frontend API hooks & visual search event handling) * `[NEW] images/` (Directory storing actual product images generated locally) --- ## Verification Plan ### Automated Verification 1. Run server `uvicorn app:app --reload` 2. Test routes with curl/postman: - Text Search: `curl -X POST "http://127.0.0.1:8000/api/search" -H "Content-Type: application/json" -d "{\"query\": \"red sneakers\"}"` - Image Search: Test image upload route using a Python script or cURL multipart upload. 3. Assert that similarity search returns expected ranked orders (e.g., searching "wallet" ranks leather wallet first). ### Manual Verification 1. Verify the visual web interface by opening it in the browser. 2. Drag and drop a sneaker image and ensure it matches sneaker images in the database. 3. Click "Find Similar" on the gold watch and ensure it returns watch-related products first. 4. Verify responsiveness and styling details on standard desktop and mobile resolutions.