coolblaze03 commited on
Commit
28ff805
·
verified ·
1 Parent(s): f534783

docs: disclose the 19.4% ast.unparse formatting ceiling (new section 8 + section 7 table row)

Browse files
Files changed (1) hide show
  1. ORACLE-LIMITS.md +54 -0
ORACLE-LIMITS.md CHANGED
@@ -131,3 +131,57 @@ the unverified remainder as errors understates the model; treating it as correct
131
  | `-O` mismatch collapse, docstrings unprovable | this file; `EVAL.md`; `weights/MODEL-CARD.md` |
132
  | unverified ≠ wrong | this file; `EVAL.md`; `weights/MODEL-CARD.md`; `harness/README.md` |
133
  | 3.13 / PyInstaller untested | this file; `EVAL.md`; `weights/MODEL-CARD.md` |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
131
  | `-O` mismatch collapse, docstrings unprovable | this file; `EVAL.md`; `weights/MODEL-CARD.md` |
132
  | unverified ≠ wrong | this file; `EVAL.md`; `weights/MODEL-CARD.md`; `harness/README.md` |
133
  | 3.13 / PyInstaller untested | this file; `EVAL.md`; `weights/MODEL-CARD.md` |
134
+ | **19.4% formatting ceiling (§8)** | **this file** |
135
+
136
+ ## 8. The formatting ceiling: 19.4% of real modules cannot certify a re-formatted answer
137
+
138
+ This section was added on 2026-09-11. It discloses a ceiling that was measured earlier and was
139
+ not previously written down here, which left §7's promise of a complete accounting unmet.
140
+
141
+ **`ast.unparse` is not bytecode-preserving.** Compiling a source file, and compiling that same
142
+ file after an AST round trip, do not always produce the same code object:
143
+
144
+ ```python
145
+ compile(ast.unparse(ast.parse(src))) != compile(src) # for 100 of 516 stdlib modules
146
+ ```
147
+
148
+ Measured on all 516 CPython 3.12 stdlib modules: **100 fail the round trip — an 80.6% pass rate,
149
+ so 19.4% of real modules contain at least one affected construct.**
150
+
151
+ **The cause**, reduced to a repro verified on the measurement box:
152
+
153
+ ```python
154
+ a = {name for name, value in ns.items() if getattr(value, 'x', False)} # POP_JUMP_IF_FALSE
155
+ a = {name # POP_JUMP_IF_TRUE
156
+ for name, value in ns.items() # + JUMP_BACKWARD
157
+ if getattr(value, 'x', False)}
158
+ ```
159
+
160
+ Identical AST, different bytecode. CPython 3.12's CFG optimiser lays out basic blocks according to
161
+ the **physical line layout** of a comprehension. An `if` *statement* wrapped the same way does not
162
+ change; comprehensions do.
163
+
164
+ **What this means for anyone reading a score from this oracle.**
165
+
166
+ - It is a ceiling on the **oracle**, not a defect in any model. A decompilation that is
167
+ semantically perfect but wraps a comprehension across lines differently from the original will
168
+ **not certify**. It is scored as unverified, which per §5 means *unknown*, not *wrong*.
169
+ - It applies to **every system measured against this oracle** — ours, PyLingual's, and any other.
170
+ It is not a differential advantage or disadvantage to anybody.
171
+ - It is **whole-module scale**: 19.4% is the fraction of *modules* containing at least one
172
+ affected construct, not the fraction of constructs affected. Function- and class-level units are
173
+ affected at a lower rate, because the chance of containing a comprehension is lower.
174
+ - **It compounds with §2's 0.33% floor and §3's `-O` collapse.** These are independent sources of
175
+ false rejection; none of them ever produces a false accept.
176
+
177
+ **Why the numbers already published are not invalidated by this disclosure.** Every published
178
+ score was produced by grading a model's output against a reference *source*, never against an
179
+ `ast.unparse` round trip, so no published figure was computed through the lossy path. This section
180
+ discloses a bound on how high any such score could ever go; it does not move one. It also rules
181
+ out a class of pipeline we did not build: an AST-splice reassembler would have discarded roughly
182
+ one module in five before asking the model a single question.
183
+
184
+ **Measured, and stated as such.** 100/516 is a direct count over the CPython 3.12 stdlib on the
185
+ measurement box. The corresponding ceiling on the internal 522-row module-shaped eval set — 92.5%
186
+ overall, 88.3% on rows of ≥400 representation lines — is measured on a different set and is
187
+ reported with those internal results rather than here.