--- language: - ko - en library_name: pytorch license: other pipeline_tag: feature-extraction tags: - online-handwriting - mathematical-expression-recognition - trajectory - temporal-convolution - on-device - pytorch --- # AIFlow Math Ink 0.6 AIFlow Math Ink 0.6은 수학 필기를 **이미지보다 stroke trajectory로 먼저 해석하는** PyTorch 연구 모델이다. 온라인 펜 입력을 시간 순서가 있는 point sequence로 표현하고, 각 고립 기호에 대해 다음 두 종류의 확률을 예측한다. - 378개 exact symbol label - 324개 visual family 핵심 연구 주제는 다음과 같다. 1. 서로 다른 필기 속도를 일정한 temporal representation으로 정규화하는 방법 2. 온라인 stroke와 raster 이미지를 하나의 trajectory encoder로 처리하는 방법 3. 시각적으로 유사한 기호의 형태와 의미를 분리하는 방법 4. 기호 분류와 수식 문맥을 결합하는 방법 ![AIFlow Math Ink 0.6 architecture](assets/architecture.svg) ## 연구 아이디어 ### 1. Trajectory-first representation 일반적인 이미지 분류 방식은 필기를 고정된 bitmap으로 변환하면서 stroke 순서와 시간 정보를 제거한다. AIFlow Math Ink는 필기를 다음과 같은 시계열로 표현한다. ```text point₁ → point₂ → point₃ → ... → pen-up ``` 각 point에는 위치뿐 아니라 다음 정보가 포함된다. - 이동 방향 - 곡률 - 필기 속도 - stroke 내부 진행률 - pen-up 상태 - 전체 기호 안에서의 상대 위치 - timestamp 관측 여부 이 표현은 같은 모양이라도 stroke 순서와 움직임이 다른 필기를 구분할 수 있게 한다. ### 2. Canonical temporal sampling 사람마다 필기 속도와 touch event 발생 빈도가 다르기 때문에 원본 event 개수는 일정하지 않다. 모델은 각 stroke를 6Hz canonical timeline으로 재표본화한다. 이때 다음 anchor는 항상 보존한다. - stroke 시작점 - stroke 끝점 - pen-up - 원본 timestamp 이를 통해 빠르게 쓴 기호와 느리게 쓴 기호를 비슷한 temporal scale에서 비교할 수 있다. ```text Raw touch events → stroke anchor preservation → 6 Hz canonical sampling → maximum 128 events ``` 원본 touch event는 그대로 유지하고, 재표본화된 sequence는 모델 입력용 view로 사용한다. ### 3. Shape coordinate와 canvas coordinate의 분리 기호의 모양과 수식 안에서의 위치는 서로 다른 정보를 가진다. AIFlow Math Ink는 두 좌표계를 동시에 사용한다. ```text shape_x, shape_y ``` 기호 자체의 bounding box를 기준으로 정규화한 좌표다. 글자의 크기와 위치 영향을 줄이고 순수한 형태를 표현한다. ```text canvas_x, canvas_y ``` 전체 canvas 또는 수식 영역을 기준으로 한 좌표다. 기호의 상대 크기와 수직 위치를 보존한다. 예를 들어 소문자 `x`, 대문자 `X`, 곱셈 기호 `×`는 모양이 비슷하지만 수식 안에서의 크기와 위치 분포가 다를 수 있다. 두 좌표계를 함께 사용하면 형태와 문맥을 분리해 학습할 수 있다. ### 4. Online과 raster의 공통 trajectory space 온라인 입력에는 stroke sequence가 직접 존재하지만, raster 이미지에는 stroke 순서가 없다. Raster 경로는 이미지를 곧바로 symbol label로 분류하는 대신, 이미지에서 여러 개의 virtual trajectory를 생성한다. ```text Raster image → spatial encoder → virtual-trajectory decoder → top-4 stroke hypotheses → shared trajectory encoder ``` 각 hypothesis는 다음 정보를 포함한다. - virtual stroke coordinates - stroke state - progress - hypothesis score 이후 온라인 입력과 동일한 residual TCN이 virtual trajectory를 처리한다. 이 구조의 목적은 입력 modality가 달라도 최종 인식은 동일한 trajectory representation 위에서 수행하는 것이다. ### 5. Exact label과 visual family의 계층적 예측 수학 기호에는 형태가 거의 동일하지만 의미가 다른 경우가 많다. 예: ```text O / 0 / o x / X / × | / 1 / l ``` 모델은 이를 하나의 분류 문제로만 다루지 않고 두 개의 prediction head로 나눈다. ```text Shared trajectory representation ├─ exact head: 378 labels └─ family head: 324 visual families ``` Exact head는 최종 symbol token을 예측한다. Family head는 시각적으로 가까운 기호 집합을 예측한다. 이를 통해 trajectory encoder가 형태를 먼저 안정적으로 구분하고, 정확한 의미 선택은 별도의 문맥 정보와 결합할 수 있다. ### 6. Formula-context behavior modeling 고립 기호의 trajectory만으로 구분하기 어려운 기호에는 수식 문맥을 추가한다. Behavior head는 다음 두 표현을 결합한다. ```text 128×19 stroke representation + 49-dimensional formula context ``` Context feature에는 다음과 같은 정보가 포함된다. - 수식 전체 크기 대비 기호 크기 - baseline과의 상대 위치 - 주변 기호의 위치와 간격 - symbol group geometry - 이웃 token 패턴 현재 behavior head의 대표 역할 분류는 다음과 같다. ```text identifier_lower identifier_upper multiply_operator ``` 예를 들어 `x`, `X`, `×`의 trajectory가 비슷할 때, 기호의 상대 크기와 주변 token을 이용해 역할을 선택한다. ### 7. Boundary behavior modeling 수식 인식에서는 여러 stroke를 하나의 기호로 묶는 grouping 과정이 필요하다. Boundary behavior head는 segmentation candidate가 실제 기호 경계를 가로지르는지를 geometry feature로 예측한다. ```text Segmentation lattice geometry → 17-dimensional boundary features → boundary behavior score ``` 이 score는 가까이 있는 두 기호가 하나의 symbol group으로 합쳐지는 overmerge를 억제하는 데 사용된다. ## 입력 표현 모델 입력은 최대 128개의 event와 19개의 feature로 구성된다. ```text input shape: 128 × 19 sample rate: 6 Hz normalized ink space: 128 × 128 ``` ![Stroke encoding](assets/stroke_encoding.svg) ### 19 input channels ```text shape_x, shape_y, canvas_x, canvas_y, direction_x, direction_y, curvature, pen_up, stroke_progress, aspect_ratio, bbox_top, bbox_bottom, bbox_height, center_y, baseline_available, time_delta, speed, missing_mask, source_modality ``` | Feature group | 역할 | |---|---| | `shape_x`, `shape_y` | 기호 내부의 정규화된 형태 | | `canvas_x`, `canvas_y` | 전체 canvas 안에서의 상대 위치 | | `direction_x`, `direction_y` | point 이동 방향 | | `curvature` | stroke의 국소 곡률 | | `pen_up` | stroke 경계 | | `stroke_progress` | stroke 시작부터 끝까지의 진행도 | | `aspect_ratio` | 기호 bounding box 비율 | | `bbox_*`, `center_y` | 수식 안에서의 geometry | | `baseline_available` | baseline feature의 유효성 | | `time_delta`, `speed` | temporal dynamics | | `missing_mask` | 관측되지 않은 시간 정보 표시 | | `source_modality` | online 또는 raster 입력 구분 | Timestamp가 존재하는 입력은 관측된 시간 차이와 속도를 사용한다. 정적 이미지에서 생성된 virtual trajectory는 canonical timing과 missingness 정보를 함께 사용한다. ## 모델 구조 ```text Online strokes → canonical sampler → 128×19 sequence → dual-TCN online adapter ────────────────────────────────────────────┐ │ Raster 128×128 │ → spatial encoder │ → causal virtual-trajectory decoder │ → top-4 trajectory hypotheses │ ────────────────────────────────────────────┤ ▼ shared residual TCN ├─ exact head └─ visual-family head ``` 주요 설정: | 항목 | 값 | |---|---:| | Maximum events | 128 | | Input features | 19 | | Shared hidden size | 128 | | Exact labels | 378 | | Visual families | 324 | | Raster hypotheses | 4 | | Canonical sample rate | 6Hz | ### Dual-TCN online adapter 온라인 입력은 두 종류의 temporal pattern을 함께 처리한다. - 기호 전체의 긴 stroke 흐름 - point 사이의 짧은 국소 변화 Dual-TCN adapter는 서로 다른 receptive field를 가진 temporal convolution을 결합해 이 두 패턴을 표현한다. Adapter의 shared state는 base trajectory encoder와 함께 합성된다. ```text base trajectory state → online shared state → modality-specific adaptation → classifier heads ``` ## Checkpoints 각 seed는 다음 세 파일로 구성된다. ```text models/ seed17/ base_378.pt online_adapter.pt behavior_role_head.pt seed31/ base_378.pt online_adapter.pt behavior_role_head.pt seed47/ base_378.pt online_adapter.pt behavior_role_head.pt ``` | 파일 | 내용 | |---|---| | `base_378.pt` | 공통 trajectory encoder, raster encoder, exact/family classifier | | `online_adapter.pt` | 실제 online stroke에 대한 dual-TCN adaptation | | `behavior_role_head.pt` | 기호의 수식 내 역할을 분류하는 context head | 세 seed는 서로 다른 초기화에서 학습되어 모델 간 오류 합의와 ensemble 특성을 분석하는 데 사용된다. ## 사용 예시 ```python from pathlib import Path from math_grid_drawer.research.math_ink_06 import MathInk06Engine root = Path("models/seed17") engine = MathInk06Engine( root / "base_378.pt", adapter_checkpoint=root / "online_adapter.pt", ) result = engine.recognize_online( strokes, canvas_width=128, canvas_height=128, top_k=5, ) ``` 출력에는 exact symbol 후보와 각 후보의 확률이 포함된다. ```python for candidate in result.candidates: print(candidate.token, candidate.probability) ``` ## 연구 결과 ### 고립 기호 trajectory classification Seed 17의 writer-disjoint 평가 결과: | Split | Exact top-1 | Exact top-5 | Family top-1 | |---|---:|---:|---:| | Writer validation, 4,261 samples | 86.13% | 99.48% | 92.94% | | Paired test, 3,782 samples | 82.87% | 97.73% | 91.22% | 세 seed의 확률을 평균한 ensemble 결과: | Metric | Result | |---|---:| | Exact top-1 | 83.71% | | Exact top-5 | 98.02% | | Visual-family top-1 | 92.99% | | Seed oracle top-1 | 88.05% | Exact top-1과 family top-1의 차이는 전체적인 형태 인식보다 동일 family 내부의 의미 선택이 더 어려운 문제임을 보여준다. Exact 오류의 56.98%는 다음과 같은 동일 visual-family 내부 혼동이었다. ```text O / 0 / o uppercase / lowercase vertical-line family cross family ``` 이 결과는 trajectory encoder가 형태 후보를 만들고, formula context가 최종 의미를 선택하는 계층적 구조를 뒷받침한다. ### Formula-context behavior head 정답 symbol grouping 위에서 평가한 3-seed 평균: | Metric | Result | |---|---:| | Role accuracy | 93.33% | | Macro-F1 | 76.41% | | Lowercase identifier recall | 95.66% | | Uppercase identifier recall | 55.91% | | Multiplication recall | 88.89% | | Expected calibration error | 4.88% | Behavior head를 teacher prediction 뒤에 적용했을 때 대상 기호의 exact top-1은 다음과 같이 변화했다. ```text 33.65% → 47.30% ``` 평균 상승 폭은 `13.65%p`였으며 rewrite precision은 `81.35%`였다. 이 실험은 trajectory만으로 선택하기 어려운 exact meaning을 수식 문맥이 보완할 수 있음을 보여준다. ### Continuous-formula adaptation Product trajectory encoder와 online adapter를 고정하고, 연속 수식의 symbol group에 작은 formula adapter를 추가했다. ```text Frozen trajectory encoder + frozen online adapter + hidden-64 formula adapter ``` 3-seed 평균 결과: | Metric | Mean | |---|---:| | Writer-validation exact top-1 | 85.19% | | Writer-validation family top-1 | 93.30% | | Official-test exact top-1 | 82.38% | | Official-test family top-1 | 88.17% | 작은 adapter만으로도 고립 기호 encoder의 representation을 연속 수식 환경에 맞게 조정할 수 있음을 확인했다. ### Grouping boundary head | Metric | Base grouping | Boundary head | |---|---:|---:| | Exact partition | 60.04% | 60.25% | | Pair-F1 | 91.07% | 91.26% | | Overmerge formula rate | 21.72% | 20.49% | Boundary head는 기호 사이의 geometry를 이용해 인접 symbol의 과도한 병합을 줄였다. ### Family fusion Validation에서 exact probability와 family probability를 결합하는 fusion weight를 선택했다. ```text exact score + 0.15 × family-consistency score ``` Paired test exact top-1: ```text 83.71% → 83.82% ``` Family prediction은 exact classifier의 후보 순서를 작은 범위에서 보정하는 역할을 한다. ## 연구 해석 AIFlow Math Ink 0.6의 실험은 수학 필기 인식을 다음 세 단계로 분리할 수 있음을 보여준다. ```text 1. Stroke trajectory에서 시각적 형태를 추출 2. Visual family 안에서 가능한 symbol 후보를 구성 3. 수식 문맥과 geometry로 exact meaning을 선택 ``` 이 구조는 고립 기호 인식과 전체 수식 해석을 하나의 거대한 classifier로 처리하는 대신, 각 단계가 서로 다른 정보를 담당하도록 한다. - Trajectory encoder: 필기 동작과 형태 - Family head: 시각적 유사성 - Behavior head: 기호의 문법적 역할 - Boundary head: stroke grouping - Formula adapter: 수식 안에서의 상대 크기와 위치 ## 연구 자료 - [전체 기술 보고서](reports/RESEARCH_REPORT.md) - [Federation audit](reports/FEDERATION_AUDIT_20260724.md) - [3-seed behavior 결과](reports/behavior_role_3seed_summary.json) - [Boundary 결과](reports/boundary_behavior_guard_report.json) - [Online 오류 합의 분석](reports/online_error_consensus_paired_test_3seed.json) - [Family-fusion calibration](reports/online_family_fusion_calibration_3seed.json) - [Formula-context 계약 분석](reports/formula_context_contract_3seed.json) - [Formula adapter 결과](reports/formula_adapter_h64_3seed_summary.json) - [Checkpoint manifest](MANIFEST.json)