Momo 1.0 β€” the decision brain of Babymomo 🧠

Momo 1.0 is a tiny, fast, fully self-contained decision model. Whatever the user says, Momo decides everything by itself β€” what to remember, what to update, what to forget, what to search in personal memory, what to search on Google, and what to anonymize before it ever leaves the device. There are no agents.md files, no skills folders, no LangChain agents β€” every decision comes from the model weights.

  • Name: Momo
  • Version: 1.0
  • Repo: momo-1.0
  • Backbone: Qwen/Qwen2-0.5B-Instruct + LoRA adapter (r=32, Ξ±=64, all attention+MLP projections)
  • Training data: data/brain.jsonl β€” 3,000 rows of decision supervision
  • Memory embeddings (optional): BAAI/bge-small-en-v1.5
  • Target latency: <100 ms memory filtering; single-decision JSON output

The decision contract

Momo receives one user message and replies with exactly one JSON object:

{
  "action": "STORE_MEMORY|UPDATE_MEMORY|DELETE_MEMORY|MEMORY_ONLY|WEB_ONLY|HYBRID",
  "memory_query": "",
  "web_query": "",
  "need_memory_search": false,
  "need_web": false,
  "memory_text": ""
}
Action Meaning Pipeline effect
STORE_MEMORY user shares a new personal fact memory_text is saved to memory
UPDATE_MEMORY correction / changed fact memory searched by memory_query, old record replaced by memory_text
DELETE_MEMORY "forget that…" memory searched by memory_query, matched records deleted
MEMORY_ONLY question about the user's own life memory searched with memory_query, answered from memory (Source of Truth)
WEB_ONLY general / world / external question Google (Serper) searched with web_query
HYBRID personal + world mix (or save + search) memory and web paths both run; may also store memory_text

Special MEMORY_ONLY with empty queries/flags = pure chit-chat (no retrieval needed).

Privacy: anonymization before the web

Memory is the Source of Truth; Google search is extra. Momo is trained to strip all private details β€” names of people in the user's life, phone numbers, emails, IDs β€” from web_query before anything is sent to Google:

"My girlfriend Sarah loves hiking β€” gift ideas for her birthday?" β†’ memory_query: "Sarah gift preferences and interests" (stays local) β†’ web_query: "birthday gift ideas for girlfriend" (anonymized)

app.py adds a second, regex-based safety net (anonymize_web_query) that scrubs any leak before the Serper call. City names are kept only when the query is location-dependent (e.g. weather), which carries no personal identity.

Repository layout

momo-1.0/
β”œβ”€β”€ data/brain.jsonl              # 3,000-row decision-training dataset
β”œβ”€β”€ train.py                      # LoRA training (T4-friendly)
β”œβ”€β”€ test.py                       # local test battery + PII-leak checks + latency
β”œβ”€β”€ app.py                        # FastAPI / Lightning App: decide β†’ search β†’ answer
β”œβ”€β”€ momo_core.py                  # single source of truth: prompt, schema, parser, validator
β”œβ”€β”€ scripts/
β”‚   β”œβ”€β”€ generate_brain_data.py    # regenerates data/brain.jsonl (seed=42)
β”‚   β”œβ”€β”€ validate_brain_data.py    # dataset schema + anonymization validator
β”‚   └── setup_git.sh              # git init β†’ GitHub repo β†’ push β†’ set HF_TOKEN secret
β”œβ”€β”€ .github/workflows/push-to-hf.yml  # mirrors every push to HF: Ansaribilal/momo-1.0
β”œβ”€β”€ requirements.txt
└── README.md

Dataset: data/brain.jsonl

3,000 rows in chat format (system β†’ compact Momo contract, user β†’ message, assistant β†’ canonical JSON decision). Distribution:

Action Rows Examples
STORE_MEMORY 700 "Remember that my favorite food is sushi."
UPDATE_MEMORY 500 "I moved from Pune to Bangalore, update my city."
DELETE_MEMORY 450 "Forget my old phone number."
MEMORY_ONLY 560 "What's my favorite food?" (+ chit-chat rows)
WEB_ONLY 460 "Is the Pixel 8 worth buying?" (+ anonymization-heavy rows)
HYBRID 330 "My boss Priya said keto is bad β€” is that true?"

Every row passed scripts/validate_brain_data.py: schema validity, flag consistency, and a hard PII-leak scan proving no private name ever appears in a web_query.

Training (Lightning AI free tier)

Runs on a T4 (16 GB) in ~10 minutes, costing well under the 15 credits/month budget. Loss is computed only on the assistant JSON (prefix masking), fp16, cosine LR, 3 epochs, batch 8 Γ— grad-accum 2, LR 2e-4.

pip install -r requirements.txt
python train.py                      # β†’ ./adapter (LoRA weights + tokenizer + meta)
python train.py --push Ansaribilal/momo-1.0   # optionally push adapter to HF

Test locally

python test.py --mock     # logic-only: parser + validator round-trips (no model)
python test.py            # full battery: 16 cases, action accuracy, PII checks, latency

Serve

export SERPER_API_KEY=...            # google.serper.dev
export HF_TOKEN=...                  # to pull the adapter from HF (or keep ./adapter local)
python app.py                        # FastAPI on :8000  (also runs in a Lightning Studio)
lightning run app app.py             # Lightning App mode
Endpoint Purpose
POST /decide fast path: text β†’ Momo JSON decision only
POST /chat full pipeline: decision β†’ memory ops β†’ memory search (bge-small) β†’ Serper β†’ final answer
GET /health status, device, avg decision latency, web enabled
curl -s localhost:8000/chat -H 'Content-Type: application/json' \
  -d '{"text":"My girlfriend Sarah loves hiking - gift ideas for her birthday?","user_id":"bilal"}'

Response includes the decision, the anonymized web_query_sent, memory hits with relevance scores, memory ops performed, and per-stage timings (memory_filter ms).

CI/CD

  • scripts/setup_git.sh β€” creates the GitHub repo, pushes, and uploads the HF_TOKEN Actions secret automatically.
  • .github/workflows/push-to-hf.yml β€” on every push to main, mirrors the entire repo (code, dataset, trained adapter) to https://huggingface.co/Ansaribilal/momo-1.0.

Limitations

  • English-only (v1.0); the decision JSON is constrained but not formally grammar-enforced β€” momo_core.parse_decision repairs minor slips and falls back to a safe memory search.
  • Memory search quality depends on the backend; LocalJSONMemory is a fallback β€” plug Babymomo's production memory backend via get_backend().
  • Momo decides and routes; open-domain answer quality comes from the backbone (0.5B) β€” intended for decision-making + retrieval grounding, not encyclopedic generation.
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Ansaribilal/momo-1.0

Base model

Qwen/Qwen2-0.5B
Adapter
(507)
this model