--- license: apache-2.0 base_model: Qwen/Qwen2-0.5B-Instruct tags: - babymomo - momo - decision-brain - memory - personal-ai - lora - qwen2 - json-constrained language: - en library_name: peft pipeline_tag: text-generation --- # Momo 1.0 β€” the decision brain of Babymomo 🧠 Momo 1.0 is a tiny, fast, fully self-contained decision model. Whatever the user says, Momo decides **everything by itself** β€” what to remember, what to update, what to forget, what to search in personal memory, what to search on Google, and what to anonymize before it ever leaves the device. There are **no agents.md files, no skills folders, no LangChain agents** β€” every decision comes from the model weights. - **Name:** Momo - **Version:** 1.0 - **Repo:** `momo-1.0` - **Backbone:** `Qwen/Qwen2-0.5B-Instruct` + LoRA adapter (r=32, Ξ±=64, all attention+MLP projections) - **Training data:** `data/brain.jsonl` β€” 3,000 rows of decision supervision - **Memory embeddings (optional):** `BAAI/bge-small-en-v1.5` - **Target latency:** <100 ms memory filtering; single-decision JSON output ## The decision contract Momo receives one user message and replies with **exactly one JSON object**: ```json { "action": "STORE_MEMORY|UPDATE_MEMORY|DELETE_MEMORY|MEMORY_ONLY|WEB_ONLY|HYBRID", "memory_query": "", "web_query": "", "need_memory_search": false, "need_web": false, "memory_text": "" } ``` | Action | Meaning | Pipeline effect | |---|---|---| | `STORE_MEMORY` | user shares a new personal fact | `memory_text` is saved to memory | | `UPDATE_MEMORY` | correction / changed fact | memory searched by `memory_query`, old record replaced by `memory_text` | | `DELETE_MEMORY` | "forget that…" | memory searched by `memory_query`, matched records deleted | | `MEMORY_ONLY` | question about the user's own life | memory searched with `memory_query`, answered from memory (Source of Truth) | | `WEB_ONLY` | general / world / external question | Google (Serper) searched with `web_query` | | `HYBRID` | personal + world mix (or save + search) | memory **and** web paths both run; may also store `memory_text` | Special `MEMORY_ONLY` with empty queries/flags = pure chit-chat (no retrieval needed). ## Privacy: anonymization before the web Memory is the **Source of Truth**; Google search is *extra*. Momo is trained to strip all private details β€” names of people in the user's life, phone numbers, emails, IDs β€” from `web_query` before anything is sent to Google: > "My girlfriend Sarah loves hiking β€” gift ideas for her birthday?" > β†’ `memory_query: "Sarah gift preferences and interests"` (stays local) > β†’ `web_query: "birthday gift ideas for girlfriend"` (anonymized) `app.py` adds a second, regex-based safety net (`anonymize_web_query`) that scrubs any leak before the Serper call. City names are kept only when the query is location-dependent (e.g. weather), which carries no personal identity. ## Repository layout ``` momo-1.0/ β”œβ”€β”€ data/brain.jsonl # 3,000-row decision-training dataset β”œβ”€β”€ train.py # LoRA training (T4-friendly) β”œβ”€β”€ test.py # local test battery + PII-leak checks + latency β”œβ”€β”€ app.py # FastAPI / Lightning App: decide β†’ search β†’ answer β”œβ”€β”€ momo_core.py # single source of truth: prompt, schema, parser, validator β”œβ”€β”€ scripts/ β”‚ β”œβ”€β”€ generate_brain_data.py # regenerates data/brain.jsonl (seed=42) β”‚ β”œβ”€β”€ validate_brain_data.py # dataset schema + anonymization validator β”‚ └── setup_git.sh # git init β†’ GitHub repo β†’ push β†’ set HF_TOKEN secret β”œβ”€β”€ .github/workflows/push-to-hf.yml # mirrors every push to HF: Ansaribilal/momo-1.0 β”œβ”€β”€ requirements.txt └── README.md ``` ## Dataset: `data/brain.jsonl` 3,000 rows in chat format (`system` β†’ compact Momo contract, `user` β†’ message, `assistant` β†’ canonical JSON decision). Distribution: | Action | Rows | Examples | |---|---|---| | `STORE_MEMORY` | 700 | "Remember that my favorite food is sushi." | | `UPDATE_MEMORY` | 500 | "I moved from Pune to Bangalore, update my city." | | `DELETE_MEMORY` | 450 | "Forget my old phone number." | | `MEMORY_ONLY` | 560 | "What's my favorite food?" (+ chit-chat rows) | | `WEB_ONLY` | 460 | "Is the Pixel 8 worth buying?" (+ anonymization-heavy rows) | | `HYBRID` | 330 | "My boss Priya said keto is bad β€” is that true?" | Every row passed `scripts/validate_brain_data.py`: schema validity, flag consistency, and a hard **PII-leak scan** proving no private name ever appears in a `web_query`. ## Training (Lightning AI free tier) Runs on a **T4 (16 GB)** in ~10 minutes, costing well under the 15 credits/month budget. Loss is computed **only on the assistant JSON** (prefix masking), fp16, cosine LR, 3 epochs, batch 8 Γ— grad-accum 2, LR 2e-4. ```bash pip install -r requirements.txt python train.py # β†’ ./adapter (LoRA weights + tokenizer + meta) python train.py --push Ansaribilal/momo-1.0 # optionally push adapter to HF ``` ## Test locally ```bash python test.py --mock # logic-only: parser + validator round-trips (no model) python test.py # full battery: 16 cases, action accuracy, PII checks, latency ``` ## Serve ```bash export SERPER_API_KEY=... # google.serper.dev export HF_TOKEN=... # to pull the adapter from HF (or keep ./adapter local) python app.py # FastAPI on :8000 (also runs in a Lightning Studio) lightning run app app.py # Lightning App mode ``` | Endpoint | Purpose | |---|---| | `POST /decide` | fast path: text β†’ Momo JSON decision only | | `POST /chat` | full pipeline: decision β†’ memory ops β†’ memory search (bge-small) β†’ Serper β†’ final answer | | `GET /health` | status, device, avg decision latency, web enabled | ```bash curl -s localhost:8000/chat -H 'Content-Type: application/json' \ -d '{"text":"My girlfriend Sarah loves hiking - gift ideas for her birthday?","user_id":"bilal"}' ``` Response includes the decision, the **anonymized** `web_query_sent`, memory hits with relevance scores, memory ops performed, and per-stage timings (`memory_filter` ms). ## CI/CD - `scripts/setup_git.sh` β€” creates the GitHub repo, pushes, and uploads the `HF_TOKEN` Actions secret automatically. - `.github/workflows/push-to-hf.yml` β€” on every push to `main`, mirrors the entire repo (code, dataset, trained adapter) to **https://huggingface.co/Ansaribilal/momo-1.0**. ## Limitations - English-only (v1.0); the decision JSON is constrained but not formally grammar-enforced β€” `momo_core.parse_decision` repairs minor slips and falls back to a safe memory search. - Memory search quality depends on the backend; `LocalJSONMemory` is a fallback β€” plug Babymomo's production memory backend via `get_backend()`. - Momo decides and routes; open-domain answer quality comes from the backbone (0.5B) β€” intended for decision-making + retrieval grounding, not encyclopedic generation.