ContinousAstra_v1 / AEGIS.md
AGofficial's picture
Upload 21 files
7bb8aac verified
|
Raw History Blame Contribute Delete
7.21 kB
# Aegis 0.1 reference runtime
`.aeg` programs run as a closed AST over integers, finite floats, booleans and strings.
The runtime uses a dedicated lexer/parser, never Python `eval`, `exec`, imports from
source, host object traversal, or shell execution. It parses and validates the entire
program, including unreachable branches, before requesting input or emitting output.
## Agent tools
Call `aegis_docs()` before writing scripts. It returns language rules and six complete
correct example `.aeg` files, each with its content, purpose, sample input strings and
expected output. Examples cover arithmetic, string concatenation, branching, nested
conditions, short-circuiting and single/multiple input pauses. These are bundled
documentation, independent of the virtual store; the tool takes no filename and reads
no host files. Create an example as a virtual file before running it.
1. `create_file(filename="scripts/greet.aeg", content="...")` creates a reusable virtual script.
2. `run_ag_file(filename="scripts/greet.aeg")` starts a fresh execution.
3. If its status is `input_required`, call `provide_ag_input(value="text")` on the next
response. The request includes a one-based index and source line/column. Read the
script and its captured output to determine what text to provide; `input()` has no prompt.
4. Continue until `ok` or `error`. Output is cumulative, including across input pauses.
Errors return empty output; output already delivered before a pause cannot be retracted.
Only one program can wait for input. Source is fixed for that run. Completed runs
release their environment. Every later invocation starts from scratch. Paused runs
are discarded at wake end, including on failure or reaching the response cap. The
virtual source file persists and can be run again after a restart.
An input request changes the current wake's limit from five to **20 total model
responses**, not 20 additional responses and not 20 responses per program. Even an
input request in response five enables response six. The cap resets to five next wake.
Scripts without input do not extend the wake. Tool calls made in response 20 still run,
but cannot cause response 21.
## Syntax and semantics
- UTF-8, one statement per line, optional final newline, LF or CRLF. Tabs are forbidden.
Indent each level by consistently two or four spaces. Blank/comment-only lines do
not affect indentation. Blocks must contain a statement.
- Names match `[a-z][a-z0-9_]*`. Keywords and `input`/`print` are reserved.
- Strings use double quotes and JSON escapes (`\n`, `\t`, `\"`, `\\`, `\uXXXX`, etc.).
Literal Unicode is allowed; invalid Unicode and raw control characters are rejected.
- Integers use decimal digits. Floats use a fractional part and/or exponent, such as
`3.5`, `1e3`, `1.5e-2`. Leading-dot and trailing-dot floats are not supported.
- Statements: `name = expression`, `print(expression)`, and `if expression:` with an
indented body and optional matching `else:` body. Assignment has program scope.
- `input()` is a string expression and can appear anywhere an expression is allowed.
`print` is a statement and accepts exactly one value. It emits a newline and spells
booleans `true`/`false`. Text is captured, never sent to a terminal.
- Precedence, high to low: parentheses; unary `not`/`-`; right-associative `^`;
`*`/`/`; `+`/`-`; comparisons; `and`; `or`. As specified, `-2 ^ 2` is `4`.
Consecutive unary operators require parentheses. Comparison chaining is rejected.
- Numeric operations exclude booleans. Mixed arithmetic promotes int to float;
overflow during promotion is an error. Division returns float. Negative integer
powers return float; nonnegative integer powers return int. `0 ^ 0` is `1`.
Zero division, complex results and non-finite values are errors.
- `+` also concatenates two strings. No string multiplication or implicit conversion.
- Equality requires identical types: `1 == 1.0` and `true == 1` are false. Ordering
accepts same types and compatible int/float pairs; all other mixed ordering fails.
Strings compare by Unicode code point; bool ordering is false before true.
- Boolean operators require bool operands and short-circuit evaluation. Their skipped
expressions are still statically validated, but do not request input or calculate.
- Validation tracks types and definite assignment across both branches. A variable
assigned in just one branch is unavailable afterward unless assigned beforehand.
Branch type unions are allowed only where every possible operand type is valid.
Sequential reassignment may change type.
## Default bounds
Host code can supply an immutable `aegis.Limits` to `Session` or `run`; source code
cannot change limits. These bounds apply across parsing, validation and evaluation:
| Resource | Default |
| --- | ---: |
| UTF-8 source bytes | 65,536 |
| Characters per line | 2,048 |
| Identifier characters | 64 |
| Block/expression nesting | 32 |
| Tokens, including structural tokens | 8,192 |
| Budgeted operations | 20,000 |
| Conservative cumulative allocation accounting | 2,000,000 bytes |
| String / per-input UTF-8 bytes | 16,384 / 16,384 |
| Captured output UTF-8 bytes | 65,536 |
| Integer magnitude | 4,096 bits |
| Absolute exponent | 4,096 |
| Input requests per run | 20 |
| Active parse/validate/execute time | 1 second |
Memory accounting charges source storage, tokens, AST nodes, values, assignments and
output. It is cumulative, so a script can exhaust it even if its current live values
would fit. Integer multiplication/power and string concatenation check growth before
performing the operation. A conservative growth check can reject a borderline value.
Nesting and token limits bound parser/validator recursion and AST growth. Time checks
occur at bounded operations; time spent waiting for the AI is excluded.
## Security boundary and its limits
The interpreter receives only source and explicit input strings. It exposes no host
filesystem, networking, process, environment, clock, random, reflection or object APIs.
Filenames are virtual dictionary keys, never paths resolved against the machine.
Outputs come only from Aegis values; unexpected failures become fixed `host_error`
messages. Structured diagnostics have a category, source line and column and contain
no host exception text, traceback, terminal capture or real source filename.
This is an application-level capability boundary, **not an absolute OS sandbox**.
The Python interpreter and trusted host runtime remain part of the trusted computing
base. Allocation accounting does not impose a process RSS limit, and timeout checks
are cooperative rather than an OS kill timer. No implementation can promise absence
of all bugs. Aegis can echo a path or secret explicitly supplied in its source/input;
it has no mechanism to discover either from the host. The agent's separately authorized
email and web tools are outside this language boundary.
Error categories: `syntax_error`, `indentation_error`, `name_error`, `type_error`,
`value_error`, `limit_error`, `host_error`. File lookup failures and malformed tool
requests use the same safe envelopes. Normal results contain only JSON data.