File size: 5,824 Bytes
73b69a3
50b69d7
 
 
 
73b69a3
 
 
50b69d7
 
 
73b69a3
 
50b69d7
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
---
title: Laya Browser Agent MCP
emoji: 🌐
colorFrom: blue
colorTo: indigo
sdk: gradio
sdk_version: 6.29.0
app_file: app.py
python_version: "3.12"
short_description: Laya browser agent decision MCP server on ZeroGPU
startup_duration_timeout: 30m
---

# 🌐 Laya Browser Agent MCP Server

A Hugging Face Gradio Space exposing a high-speed, non-autoregressive **Laya** browser decision agent as an **MCP (Model Context Protocol)** server with **ZeroGPU** acceleration and a persistent **Storage Bucket** volume (`/data`).

This Space allows any MCP client (such as Claude Desktop, Cursor, Cline, or custom agent scripts) to control a headless Playwright browser using both direct action primitives (`open_url`, `observe`, `click`, `type_text`, `select_option`, `scroll`, `screenshot`) and autonomous Laya decision steps (`laya_step`, `laya_run`).

---

## ⚑ Architecture & Features

- **System-1 Non-Autoregressive Decision Engine**: Powered by [`ichenney/laya-browser-v32b`](https://huggingface.co/ichenney/laya-browser-v32b) (322M parameters, fine-tuned from `cklxx/laya-browser`), capable of single-pass multi-choice scoring across DOM elements and high-accuracy `DONE`/`BLOCKED` judgments.
- **ZeroGPU Acceleration with Seamless CPU Fallback**: Model inference runs on dynamic NVIDIA RTX 6000 Ada / Blackwell GPUs using `@spaces.GPU(duration=15)`. If GPU quota is temporarily unavailable or exhausted, inference falls back to CPU automatically without breaking the agent loop.
- **Persistent Storage Bucket (`/data`)**: Playwright authentication and browser session state (`/data/state.json`) persist across Space restarts on a private Hugging Face Storage Bucket (`LazyHuman/laya-browser-data`).
- **Gradio Native MCP Server**: Built-in Gradio 6 MCP server (`demo.launch(mcp_server=True)`), exposing tools over SSE and JSON-RPC.

---

## πŸ› οΈ MCP Tools

| Tool | Parameters | Description |
|---|---|---|
| `open_url` | `url: str` | Navigates the browser to the requested URL and returns the page title and final URL. |
| `observe` | None | Returns page title, current URL, truncated visible text (~3000 chars), and a numbered list of visible interactive controls. |
| `click` | `index: int` | Clicks the interactive control corresponding to `index` (using the same shared enumeration as `observe`). |
| `type_text` | `index: int`, `text: str` | Fills or types text into the input or editable control at `index`. |
| `select_option` | `index: int`, `value: str` | Selects a dropdown `<select>` option for the control at `index`. |
| `scroll` | `direction: str` | Scrolls the page (`"down"` or `"up"`). |
| `screenshot` | None | Captures and returns a screenshot of the current page viewport. |
| `laya_step` | `goal: str`, `execute: bool = True` | Constructs state and question schema from the live DOM, invokes Laya decision head on ZeroGPU (with CPU fallback), and executes the chosen action if `execute=True`. |
| `laya_run` | `goal: str`, `max_steps: int = 10` | Runs an autonomous decision loop up to `max_steps`, terminating early on `DONE`, `BLOCKED`, or repeated action loops. |
| `reset_browser` | None | Closes and recreates the browser context, clearing transient state. |
| `save_session` | None | Explicitly persists Playwright `storage_state` (cookies, local storage) to `/data/state.json`. |

---

## πŸ”Œ MCP Client Configuration

### MCP Endpoint
- **URL**: `https://lazyhuman-laya-browser-mcp.hf.space/gradio_api/mcp/`

### 1. Claude Desktop / Cursor / Cline (`claude_desktop_config.json` or `.cursor/mcp.json`)
```json
{
  "mcpServers": {
    "laya-browser": {
      "url": "https://lazyhuman-laya-browser-mcp.hf.space/gradio_api/mcp/",
      "headers": {
        "Authorization": "Bearer <YOUR_HF_TOKEN>"
      }
    }
  }
}
```

> **Note on ZeroGPU Quota**: Visitors who provide their Hugging Face User Access Token via the `Authorization: Bearer <HF_TOKEN>` header consume their personal ZeroGPU quota (~5 min/day for free users, 40+ min for Pro users), ensuring faster queue priority and preventing shared quota exhaustion.

### 2. Standard I/O Bridge (via `mcp-remote`)
For MCP clients that only support `stdio` process spawning:
```json
{
  "mcpServers": {
    "laya-browser": {
      "command": "npx",
      "args": [
        "-y",
        "mcp-remote",
        "https://lazyhuman-laya-browser-mcp.hf.space/gradio_api/mcp/",
        "--header",
        "Authorization=Bearer <YOUR_HF_TOKEN>"
      ]
    }
  }
}
```

---

## πŸ”’ Authentication & Privacy

- **Optional API Key**: Set the secret `MCP_API_KEY` in Space Settings β†’ Secrets. If configured, calls require the `X-API-Key` or `Authorization: Bearer <KEY>` header matching `MCP_API_KEY`. If unset, the server is open.
- **Space Visibility**: If you set the Space to Private or Protected, all HTTP and MCP requests must provide a valid Hugging Face Access Token with read access.

---

## ⚠️ Known Limitations & Best Practices

1. **Free Spaces Sleep**: On the Hugging Face free tier, Spaces sleep after 48 hours of inactivity. The first request after sleep will take ~30-60 seconds while the container and Chromium spin up.
2. **Single Shared Session**: The Space hosts one browser context guarded by an `asyncio.Lock`. Concurrent requests from multiple clients will queue.
3. **Datacenter IP Blocking**: Cloud hosting IPs (such as Hugging Face / AWS) are flagged by aggressive Cloudflare/Akamai bot protection on certain commercial websites.
4. **Laya's Modest Capacity (322M params)**: Laya excels at fast, calibrated System-1 target selection and form filling on clear pages, but is not an autoregressive multi-step reasoning LLM. Use it as a rapid first-pass decider, and allow your primary reasoning LLM (Claude, GPT, etc.) to use `observe`, `click`, and `type_text` directly whenever a nuanced fallback is required.