File size: 12,542 Bytes
387b07d
 
 
 
 
 
8e51c76
387b07d
 
 
f104717
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
---
title: Openenv Claw1
emoji: 🚀
colorFrom: indigo
colorTo: gray
sdk: docker
app_port: 7860
pinned: false
---

# Last-Mile Delivery Optimization (OpenEnv)

OpenEnv-compatible benchmark environment for city-scale last-mile dispatch under realistic operating constraints.

## Motivation

Urban delivery systems face a difficult planning problem: dispatchers must complete as many deliveries as possible while navigating congestion, road blockages, and battery limits. Poor routing or action sequencing creates delays, SLA violations for high-priority orders, and unnecessary energy cost.

This environment models that real-world challenge in a deterministic grid world so policies can be compared fairly across difficulty levels.

## What This Environment Simulates

- Multi-order pickup and delivery workflow.
- Static obstacles, dynamic obstacles, and traffic penalties.
- Priority-aware delivery pressure (especially in hard mode).
- Battery constraints and charging behavior in hard mode.
- Deterministic seeded task generation for reproducible evaluation.

## Architecture Diagram

```mermaid
flowchart LR
	U[OpenEnv Evaluator or User] -->|HTTP| API[FastAPI API Layer\napp/main.py + app/routes.py]

	API -->|reset or step or state| ENV[LastMileDeliveryEnvironment\nenv/environment.py]
	ENV --> SIM[DeliverySimulator\nenv/simulator.py]
	SIM -->|observation, reward, done, info| API

	API -->|baseline endpoint rollout| BASE[BaselineGreedyAgent\nbaseline/baseline_agent.py]
	BASE --> ENV

	API -->|grader endpoint| GRADER[DeliveryEpisodeGrader\ngrader/grader.py]
	GRADER --> REPORT[GradeReport\nscore + metrics + safeguards]

	TASKS[tasks/easy.py + tasks/medium.py + tasks/hard.py\n+ tasks/registry.py] -->|task config| ENV
	TASKS -->|success_condition| GRADER
	TASKS -->|metadata| API
```

Flow summary:

- The API orchestrates environment state transitions and stores trajectory steps.
- The baseline endpoint runs a fresh rollout with the built-in greedy agent.
- The grader endpoint scores the recorded trajectory against task success conditions.
- Task definitions provide deterministic configuration and grading targets.

## Action Space

The action payload is a strict JSON object with exactly four fields.

| Field | Type | Required | Allowed Values | Notes |
|---|---|---|---|---|
| move | string or null | Yes | up, down, left, right, stay, null | Movement action; stay maps to no movement |
| accept_order | integer or null | Yes | order index or null | When set to n, maps to internal order_n |
| deliver_order | boolean | Yes | true or false | Deliver current order when true |
| wait | boolean | Yes | true or false | Explicit wait action when true |

Validation is strict. Exactly one action intent must be active per step.

## Observation Space

Each step returns an observation object with agent, order, and map state.

| Field | Type | Description |
|---|---|---|
| grid_width, grid_height | integer | Map dimensions |
| agent_location | object {x, y} | Current agent position |
| pending_orders | array of Order | Orders not yet delivered |
| current_order | Order or null | Active accepted/picked-up order |
| obstacles | array of Position | Static + dynamic blocked cells |
| dynamic_obstacles | array of Position | Current dynamic blocked cells |
| traffic_zones | array of TrafficZone | Cells with extra movement cost |
| charging_stations | array of Position | Recharge cells |
| battery_level | integer or null | Active in battery-enabled tasks |
| step_count | integer | Elapsed step count |
| total_reward | float | Cumulative episode reward |

## Tasks

All tasks are deterministic and expose explicit success conditions.

### Easy

- Description: simple single-order delivery on a compact map.
- Configuration: 6x6 grid, 1 order, max_steps=60.
- Constraints: no obstacles, no traffic, no battery, no priority mix.
- Success condition: completion_rate=1.0, max_steps=60, invalid_action_rate_max=0.12.

### Medium

- Description: multi-order delivery with static obstacles and traffic friction.
- Configuration: 10x10 grid, 3 to 4 orders, max_steps=140.
- Constraints: static obstacles + traffic enabled.
- Success condition: completion_rate_min=0.85, max_steps=140, invalid_action_rate_max=0.08.

### Hard

- Description: high-pressure dispatch with tight battery and SLA constraints.
- Configuration: 16x16 grid, 6 to 8 orders, max_steps=165.
- Constraints: static + dynamic obstacles, traffic, high-priority orders, multi-stop delivery, battery + recharge, delay pressure.
- Success condition: completion_rate_min=0.90, high_priority_on_time_rate_min=0.88, battery_depletion=false, max_steps=165, invalid_action_rate_max=0.04.

## Reward Design

The reward is dense and combines task completion signals with operational efficiency:

- Base step penalty: -1.0 each step.
- Delivery reward: +50.0 on successful delivery.
- Destination milestone: +10.0 when pickup/drop destination is reached.
- Invalid action penalty: -20.0.
- Traffic movement penalty: -2.0 when entering traffic cells.
- Progress shaping: signed distance-based reward clipped by configured bounds.
- Priority delivery bonus: +15.0 (high) or +5.0 (low).
- Recharge reward: +2.0 when battery is restored at charging stations.
- Delay penalty: scaled negative penalty for overdue active orders.
- Battery failure penalty: -30.0 when battery depletes to zero.

Design goal: favor correct and efficient order handling while discouraging invalid and wasteful behavior.

## Grader Logic

The grader is deterministic and aligned to each task's success_condition fields.

Main metrics:

- completion_rate: delivered_orders / total_orders.
- high_priority_on_time_rate: on-time high-priority deliveries.
- efficiency_ratio: normalized by max_steps target.
- invalid_action_rate: invalid_actions / steps_taken.

Scoring components:

- Completion component: highest weight.
- Priority SLA component: enabled when high_priority_on_time_rate_min exists.
- Efficiency component: step-budget performance.
- Penalty component: reduces score when invalid action rate exceeds target.

Edge-case safeguards include:

- Zero-step and zero-order episode handling.
- Invalid success-condition value fallback and clamping.
- Division-safe invalid-action and efficiency calculations.
- Optional battery depletion penalty when the task requires no depletion.

## Setup

### Local

1. Create and activate a virtual environment.
2. Install dependencies.
3. Run the API server on port 7860.

```bash
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
uvicorn app.main:app --host 0.0.0.0 --port 7860
```

### Docker

```bash
docker build -t delivery-openenv:latest .
docker run --rm -p 7860:7860 \
	-e HF_TOKEN="example-provider-key" \
	-e MODEL_NAME="gpt-4o-mini" \
	-e API_BASE_URL="https://api.openai.com/v1" \
	-e OPENAI_USE_MODEL="1" \
	delivery-openenv:latest
```

## API Usage

Base URL:

```text
http://0.0.0.0:7860
```

### Health

```bash
curl -s http://0.0.0.0:7860/health
```

### Frontend Results Dashboard

Open the interactive dashboard in your browser:

```bash
open http://0.0.0.0:7860/ui
```

The dashboard lets you run `reset`, `step`, `state`, `grader`, and `baseline` calls and inspect live payloads and metrics.

### List Tasks and Action Schema

```bash
curl -s http://0.0.0.0:7860/tasks
```

### Reset Environment

```bash
curl -s -X POST http://0.0.0.0:7860/reset \
	-H "Content-Type: application/json" \
	-d '{"task":"easy","seed":101}'
```

### Reset From Real-World Scenario

Use this endpoint when you want to inject externally curated map and order state instead of task-generated layouts.

```bash
curl -s -X POST http://0.0.0.0:7860/reset_from_scenario \
	-H "Content-Type: application/json" \
	-d '{
	  "seed": 99,
	  "success_condition": {
	    "completion_rate_min": 1.0,
	    "max_steps": 20,
	    "invalid_action_rate_max": 0.20,
	    "battery_depletion": false
	  },
	  "scenario": {
	    "width": 6,
	    "height": 6,
	    "max_steps": 20,
	    "agent_start": {"x": 0, "y": 0},
	    "orders": [
	      {
	        "order_id": "r1",
	        "pickup": {"x": 1, "y": 0},
	        "dropoff": {"x": 2, "y": 0},
	        "delivery_locations": [{"x": 2, "y": 0}],
	        "priority": "high",
	        "created_step": 0
	      }
	    ],
	    "obstacles": [{"x": 4, "y": 4}],
	    "dynamic_obstacles": [{"x": 4, "y": 5}],
	    "traffic_zones": [{"location": {"x": 3, "y": 0}, "extra_cost": 3}],
	    "charging_stations": [{"x": 0, "y": 0}],
	    "battery_profile": {
	      "enabled": true,
	      "capacity": 10,
	      "recharge_rate": 4,
	      "initial_level": 7
	    }
	  }
	}'
```

Contract summary for `scenario`:

- `width`, `height`, `max_steps`: grid and episode budget.
- `agent_start`: starting position.
- `orders`: list of orders with pickup/dropoff/delivery locations and priority.
- `obstacles`, `dynamic_obstacles`: blocked cells.
- `traffic_zones`: movement-cost cells (`extra_cost >= 1`).
- `charging_stations`: recharge cells.
- `battery_profile`: battery enablement, capacity, recharge rate, and optional initial level.
- `success_condition` (top-level): optional custom grader thresholds for this scenario.

### Step

```bash
curl -s -X POST http://0.0.0.0:7860/step \
	-H "Content-Type: application/json" \
	-d '{"move":null,"accept_order":null,"deliver_order":false,"wait":true}'
```

### State

```bash
curl -s http://0.0.0.0:7860/state
```

### Grade Current Episode

```bash
curl -s http://0.0.0.0:7860/grader
```

### Baseline Rollout Evaluation

```bash
curl -s "http://0.0.0.0:7860/baseline?task=easy"
```

## Baseline Scores (Reproducible Inference Run)

Measured with `inference.py --all-tasks --seed 42 --json-summary` using deterministic fallback mode (`OPENAI_USE_MODEL=0`).

| Task | Steps | Success | Score |
|---|---:|---:|---:|
| easy | 6 | true | 0.9700 |
| medium | 140 | false | 0.0001 |
| hard | 55 | false | 0.0667 |

Average score (seed=42): **0.3456**

## Real-World Validation

We tested the environment under realistic scenarios:

- Priority-based delivery selection.
- Traffic-aware routing decisions.
- Battery-constrained navigation.

Results indicate that evaluated agents can exhibit behavior aligned with real-world logistics systems under these constraints.

## OpenEnv Manifest

Project metadata and runtime wiring are defined in openenv.yaml, including:

- runtime entrypoints,
- environment class,
- task registry loader,
- grader class,
- action and observation schema declarations.

## Baseline Inference Script

The baseline runner uses the OpenAI Python client and supports the required submission variables:

- `HF_TOKEN`: API key/token used as OpenAI client `api_key`.
- `MODEL_NAME`: model identifier used for inference requests.
- `API_BASE_URL`: optional OpenAI-compatible endpoint.

Backward-compatible aliases are also supported: `OPENAI_API_KEY`, `OPENAI_MODEL`, and `OPENAI_BASE_URL`.

```bash
export HF_TOKEN="your-api-key"
export MODEL_NAME="gpt-4o-mini"
export API_BASE_URL="https://api.openai.com/v1"
export OPENAI_USE_MODEL="1"

# Reproducible baseline score across easy/medium/hard tasks.
python inference.py --all-tasks --seed 42 --json-summary
```

Notes:

- `OPENAI_USE_MODEL=1` enables remote model calls.
- `OPENAI_USE_MODEL=0` keeps deterministic heuristic fallback while still validating OpenAI client wiring.
- `--seed` controls reproducibility for baseline comparisons.

## Trainable Agent (Q-Learning)

This repository now includes a trainable tabular Q-learning agent for interactive learning on the OpenEnv loop.

Training script:

```bash
python scripts/train_q_agent.py --task easy --episodes 2000 --eval-episodes 200 --seed 42 --output models/q_agent_easy.json
```

What it does:

- Trains through repeated `reset()` / `step()` episodes.
- Learns action values for a compact state representation.
- Evaluates the learned policy and prints success/completion metrics.
- Saves a reusable model artifact (`models/q_agent_easy.json`).

Learner runtime components:

- Agent implementation: `baseline/trained_q_agent.py`
- Trainer entrypoint: `scripts/train_q_agent.py`

## Validation Commands

```bash
./scripts/validate.sh
.venv/bin/openenv validate
python -m pytest -q
```

## Test Command

```bash
python -m pytest -q
```

## Hugging Face Spaces (Docker)

1. Create a Docker Space.
2. Push this repository.
3. Set variables/secrets: HF_TOKEN, MODEL_NAME, API_BASE_URL, OPENAI_USE_MODEL.
4. Ensure exposed port is 7860.
5. Rebuild and verify /health.

Additional deployment notes are available in HF_SPACES_DEPLOYMENT.md.