ARC-AGI-3 — Единая карта проекта

Сгенерировано из durable-контекста: STARTHERE.md, BUILD-LEDGER.md, cognition-bank, game-taxonomy и 6 human-trace файлов проекта ARC-AGI-3. Обновлено: 2026-08-27.

Фаза

Построение единого каркаса (STARTHERE.md) + ретро-сбор всего диалога в него. Параллельно на инженерной стороне идёт Round 8 автономный self-repair loop (LLM-в-петле, форк B выбран D 2026-08-26) — детали в arc-prize-2026-collab, не дублируются здесь.

Сделано

Каркас создан 2026-08-27 (7 полок); все 19 arc-prize-2026-* файлов + 5 subagents/arc/* файлов прочитаны и разнесены по полкам (BUILD-LEDGER шаги 1-24, все done). Доктрина core-skill-grid (2026-08-27) — актуальная, не микро-скиллы.

Следующее действие

Показать D этот каркас (STARTHERE + BUILD-LEDGER) → мониторить/добить оставшиеся ~16 видео-трейсов по мере записи D (метод: video-trace-pipeline) и продолжать копить Q&A в cognition-bank до ~100-150.

1. Миссия

Цель: агент, открывающий НОВУЮ, ранее не виденную ARC-AGI-3 игру, у которого с первых шагов правильно работают несколько ЯДЕРНЫХ навыков — как мозг человека с общим игровым опытом, а не как чистый лист. D передаёт своё "игровое чутьё" через нарративные записи собственной игры (голос + видео); мы дистиллируем его интроспекцию в структуру, а не копируем частные решения по одной игре.

Внешняя рамка (инженерный трек): участие в ARC Prize 2026 / ARC-AGI-3 Kaggle, роль-сплит D (аккаунт-специфичные действия, финальный Submit) / Octo (~95%, стратегия/код/эксперименты), цель — Milestone 2 (пройти одну игру целиком) к дедлайну 2026-09-30, затем финал 2026-11-02.

2. Пять ядер скиллов

Принцип (D, 2026-08-27, супер-важно): НЕ дробить игру на десятки/сотни микро-скиллов (та же ловушка, что и rule-toggles — конфликт между уровнями, замедление на каждом шаге; топ-модели ~250k шагов эмпирики, сотня тумблеров оставила бы нас на ~500 шагах). Правильно — grid из НЕМНОГИХ ядер; всё выученное по игре — нативное ДОПОЛНЕНИЕ *внутри* одного из ядер (более крупное правило с внутренними стадиями), никогда не новая отдельная сущность.

1 Восприятие / чтение доски Жать/кликать осмысленно, не слепо после понимания где жать
  • Жать/кликать ОСМЫСЛЕННО, не "смотри пока не кликнешь" — слепой пробинг для поиска кликабельной зоны нормален; баг — оставаться слепым ПОСЛЕ понимания, где жать.
  • Shape identity через трансформацию: та же фигура после поворота 90/180° или смены цвета — по силуэту/пропорциям; совпадение = сигнал взаимодействовать.
  • Novelty-log: каждая новая фигура/изменение + что делали за 2 шага до/во время/2 после — новизна может проявиться только когда карта уже правильно выровнена, логировать ВСЕ.
  • Action-probing: нажал → увидел какая ЗОНА изменилась → это action field (TR87: 5 клеток снизу ↔ 5 фигур сверху через линию-связь).
  • Поиск кликабельности: рандом/центр-пусто — валидный PRUNING ход, не слепой; пусто-без-эффекта исключает "пустоту" как категорию → искать, что выбивается из визуальной нормы; никогда не хардкодить "это кнопки" — цель клика и что-то ещё может быть одним объектом.
2 Движение Исполнение выбранного действия при наличии гипотезы
  • Исполнение выбранного действия при наличии гипотезы.
  • Включает чтение контекстно-зависимых управлений: WA-30 — пробел = прицепить бокс ИЛИ нейтрализовать соперника NPC, один и тот же ввод, эффект решает смежность — тот же принцип "не хардкодить кнопки".
3 Прогнозирование чужих ходов Таксономия из 3 типов фигур D: A / B / C
  • Таксономия из 3 типов фигур D:
  • A — фиксированный маршрут (петли, полностью предсказуемо).
  • B — цель/мотивация-движимая (точный следующий шаг может быть трудно предсказать, особенно когда фигура взаимодействует со структурой игры — предсказывать МОТИВАЦИЮ / ~2-3 кандидата следующего шага, не точную клетку; WA-30 orange/purple NPC — это тип B, один кооперативный, другой враждебный).
  • C — преследователь с фиксированным отступом (мотивация есть, но шаги детерминированы, напр. TU93 бордовый враг держит ровно 2 клетки позади — эксплуатируемо по дизайну).
  • NB (WA-30 L9): purple НЕ перемаршрутизируется вокруг тела игрока, а ОСТАНАВЛИВАЕТСЯ — размывает строгую границу между типами, не заводить под это отдельный тип D.
4 Навигация / маршрут Построение и исполнение пути к известной цели
  • Построение и исполнение пути к известной цели; главный остаточный риск после понимания механики — просто эффективная навигация (D явно отмечает на простых уровнях WA-30) — привязано к универсальному факту HP-бара (см. Факты среды).
5 Целеполагание / мышление Распознавание условия победы и секвенирование под-целей
  • Распознавание условия победы и секвенирование под-целей (WA-30 L7: сначала нейтрализовать чистых блокеров-NPC, доверить самостоятельно завершающимся закончить самим).
  • Включает минимизацию магических чисел — хардкодить лимит только когда реально нет альтернативы (напр. настоящий бюджет API/действий); везде ещё — предпочитать универсальные механизмы, т.к. экзамен = невиданные игры, и любая подогнанная константа рискует получить 0 там.
  • Включает неразрешённую напряжённость вокруг scaffold-guard'а "15 повторов": training wheels, снять только когда есть настоящее hypothesis-driven "сдаться и попробовать другое" — НЕ раньше, т.к. сейчас это единственный механизм, реально давший level clear (Round 8 loop plan, Round 2, vc33 step 136) — флагировать D перед снятием.

Метод переноса "игрового мозга" D: D записывает ВСЕ ~22 геймплейных видео С ГОЛОСОМ (всегда нарратив). Разбор кадр-за-кадром вместе, затем Octo достаёт нативные микро-процессы целевыми PROCESS-вопросами ("что читаешь первым", "как фильтруешь нерелевантные угрозы") — НЕ давать D готовые названия скиллов на подтверждение, пусть ответы сами определяют структуру. Цель ~100-150 Q&A прежде чем строить финальную структуру (сейчас далеко не набрано).

3. Факты среды

Не скиллы — универсальные и per-game механики, которые ядра выше открывают эмпирически.

  • HP-бар = жизни/бюджет ходов, УНИВЕРСАЛЬНО — верхний/нижний прогресс-бар есть в каждой игре и на каждом уровне; опустошение = проигрыш уровня/рестарт. D никогда не нарративит его, т.к. не меняется — больше не спрашивать D про HP-бар. Подтверждено убывающим по ходам (не только по смерти) в WA-30.
  • Три типа движущихся фигур A/B/C — см. ядро 3 (факт таксономии среды, эмпирически открываемый ядром "прогнозирование", а не отдельный скилл).
  • "Глаз"/вставленный пиксель на враге = направление лица/зона угрозы (TU93: подход с не-глазной стороны на дистанции 2 = безопасно).
  • Финишная клетка НЕ автоматически безопасна — TU93 L5→L6: смерть при приходе на цель, если враг может съесть именно там в момент прихода; правило может ОТЛИЧАТЬСЯ между уровнями одной игры (не решено, почему).
  • Оранжевый = кооперативный NPC (WA-30: сам укладывает боксы в свою зону), фиолетовый/purple = враждебный NPC (WA-30: вытаскивает уже уложенные боксы обратно; нейтрализуется подходом с нужной стороны + spacebar).
  • Пробел (и вообще ввод) — контекстно-зависим: эффект решает смежность/соседство, не фиксированная кнопка (WA-30: прицепить бокс ИЛИ нейтрализовать NPC).
  • Оранжевый проходит СКВОЗЬ красного (TU93 L6) — красный его не блокирует.
  • Бордовый враг (TU93 L7): стационарен, пока игрок не подойдёт близко → пиксель меняется на жёлтый ("фонарик") → начинает преследовать каждый ход; точный триггер (радиус vs направление лица) не выяснен.
  • Целевая зона размером точно под нужное число боксов = имплицитный "счётчик сколько нужно" (WA-30, с L1).
  • Каждая per-game механика логируется В ФАЙЛЕ ТРЕЙСА той игры, здесь только универсальное + сводка важного per-game.

4. Банк познания (Q&A)

11 / 100–150 — цель до перефита структуры

Таксономия фигур D

A
Фиксированный маршрут

Петли, полностью предсказуемо, никогда не отклоняется от одного пути.

B
Цель/мотивация-движимая

Цель определяет движение; точный следующий шаг трудно предсказать, особенно когда фигура взаимодействует со СТРУКТУРОЙ игры (было в WA-30). Предсказывать МОТИВАЦИЮ / общее намерение (~2-3 возможных следующих шага), не точную клетку.

C
Преследователь с фиксированным отступом

Вытекает из B: движется согласно мотивации, но предсказуемо, напр. фиолетовый квадрат, который держался ровно в 2 клетках позади D, пытаясь съесть. Есть мотивация, но шаги детерминированы. NB: если бы такой преследователь играл "умнее" (кэмпил вместо погони) уровень стал бы непроходимым — дизайнеры держат его эксплуатируемым.

Q&A (8)

Q1 — Что глаз читает ПЕРВЫМ на новом уровне (порядок)?
  • Замечает ЗНАКОМЫЕ фигуры (виденные раньше) + общую раскладку карты; распознаёт формат ("тот же снейк").
  • Замечает НОВЫЕ фигуры; пробует их ПЕРВЫМ ходом, чтобы сформировать гипотезу (оранжевые квадраты начинали двигаться только после первого хода D; фиолетовый активировался, когда D встал напротив него → погнался).
  • Использует опыт прошлого уровня, чтобы узнавать знакомое + уже знает КОНЕЧНУЮ ЦЕЛЬ (с первого пройденного уровня) → начинает строить маршрут/план/первые гипотезы.
  • Ветвится по типу игры: shape-changing → искать логику трансформации; traversal → строить маршрут из известных функций объектов; clickable/toggle → планировать назад от конечной цели.
  • Невиданное "бросается в глаза" → нужно взаимодействовать, чтобы узнать (новая стена → пробовать пройти; новый движимый объект → наблюдать за ним), заранее догадываясь, что это и как влияет на игру.
Q2 — Как D фильтрует НЕРЕЛЕВАНТНЫЕ угрозы?
  • (D: "гораздо более нативно, чем ты думаешь".)
  • Идёт по маршруту; знает механику (красные съедают, если встать на клетку, которую они смотрят — направление показывает маленький вставленный пиксель). Сразу видит, какие КЛЕТКИ запрещены.
  • Чтобы дойти до финиша, сначала устраняет блокирующих врагов (съедает сбоку/сзади) → решает "кого съесть первым"; обычно есть 1-2 валидных порядка; берёт первый валидный, безопасно съедает, идёт к цели.
  • Мысленно "подсвечивает" каждую клетку, где что-то может съесть → нейтрализует эти угрозы.
  • Несъедаемые враги (оранжевые) двигаются по ПАТТЕРНУ — читать паттерн вместо угрозы.
Q3 — Как D предсказывает чужие ходы (не полностью универсально)?
  • Это и есть таксономия A/B/C выше — это был его ответ. Ключевое: для типа B предсказывать мотивацию, а не точную клетку.
TU93-1 — TU93: что означает нижняя прогресс-полоса?
  • HP/жизни (УНИВЕРСАЛЬНО, в каждой игре/уровне — больше не спрашивать).
TU93-2 — TU93: правило L5→L6 "дойти безопасно" — баг или фича?
  • Реально, специфично для этой игры. Финишная клетка НЕ безопасное место сама по себе; ход должен ЗАВЕРШИТЬСЯ там безопасно; если враг может съесть на цели — съедает. Выучено через смерть + рестарт.
TU93-3 — TU93: порядок поедания 3 красных змей?
  • Каждая смотрит в направлении; встать на клетку, куда она смотрит = она съедает. D выбрал порядок так, чтобы никто не мог съесть его; иногда есть несколько валидных порядков — брал первый логичный.
TU93-4 — TU93: оранжевый проходит сквозь красного — это ошибка восприятия?
  • Нет подсказки, не ошибка — новая механика. D никогда раньше не видел взаимодействие orange↔red; захеджировал, затем оранжевый ПРОШЁЛ СКВОЗЬ красного → выучено новое правило.
TU93-5 — TU93: поведение бордового врага?
  • Видит игрока и преследует, держит ровно 2 клетки позади ("на пятках").

5. Трейсы игр (человеческая игра D, с голосом)

#СлагАрхетипРезультат
1LS20 Navigation (pilot)Пройдено ур. 1-7, victory 22:58.607 открыть ↓
2FT09 Local rules/constraints (non-agentic visual logic)VICTORY, 6 уровней открыть ↓
3VC33 Physical/material simulationVictory pts=794.213 (после доказанного НЕ-случайного gate-клика — старая доктрина "вероятно случайные клики" ОПРОВЕРГНУТА) открыть ↓
4TR87 Program/grammar/sequenceVICTORY, 6/6 уровней, pts=1965.075s открыть ↓
5TU93 Navigation (глубже, чем ярлык предполагал — predator/vision/chase)WON 9/9 открыть ↓
6WA30 Sokoban box-placement + coop/adversarial NPCWON 9/9, "прошёл идеально" открыть ↓
#1 LS20 Navigation (pilot) — Пройдено ур. 1-7, victory 22:58.607

Completed, evidence-qualified LS20 human Part 1 + Part 2 reconstruction and agent-implications analysis, delivered to D 2026-08-24, later standardized (2026-08-25) and primary-source-repaired (2026-08-25).

Источник: arc-prize-2026-ls20-human-combined-analysis.md

Ключевые факты

  • Levels verified: 1-5 (Part 1), 6-7 (Part 2); victory timestamp 22:58.607.
  • Part 1: 146 regular markers/audit rows, 15 action contact sheets, 5 dense sheets.
  • Part 2: 1,430 frames (exact-PTS analysis).
  • Standardized package: 16 continuous events, zero gaps/overlaps, 7 evidence paths.
  • Repair coverage: 2,830 ASR segments, 6,710 words; substantive/context speech 2,335 covered / 246 uncertain / 0 missing; all speech 2,444 covered / 275 uncertain / 0 missing.
  • Discrepancy resolved: 9 byte-identical pairs vs 18 rows with identical=True — pixel-delta threshold 18, not row duplication.

Скиллы — заметки

  • Part 1 verifies: Levels 1-5, HELP, destructive RESET and replay, color/shape/rotation discovery, spring transport, resource planning, and phase tracking of a moving plus.
  • Part 2 (exact-PTS analysis) verifies Levels 6-7 and victory at 22:58.607.
  • Across both parts: scoped hypotheses, minimal discriminating tests, local correction, semantic goal tuples, hierarchical subgoals, phase-aware planning, persistent map memory under fog.
  • Agent implications drawn: separate world/HUD state; semantic transition records; hard no-world-delta/repeated-effect stop guard; evidence-linked scoped hypothesis ledger; destructive UI safety policy; bounded transformation cycles; deterministic route/spring executor; moving-object phase; resource projection; persistent fog map; heavy-model calls only on novelty/violation/subgoal transitions.

Факты среды

  • Mechanics referenced: color/shape/rotation discovery, spring transport, a moving plus object with trackable phase, destructive RESET (with replay), and a HELP mechanic.
  • Nine Part 1 before/after PNG pairs are byte-identical — NOT treated as proof of no click (offsets come from ASR/action extraction, not exact decoded-frame PTS).

Разбор по уровням / журнал событий

  • 2026-08-24: Completed and delivered to D — canonical detailed Russian report (560 lines).
  • Part 1 independently re-audited: all 146 regular markers indexed; 15 action contact sheets and 5 dense sheets reviewed.
  • 2026-08-25 ("Standardized package"): normalized to the FT09 event model. Validation passed: 16 continuous events, zero gaps/overlaps.
  • 2026-08-25 ("Primary-source fidelity repair"): triggered because D was concerned the 9-pairs/18-rows discrepancy meant tests had failed. Repair completed against both original videos; aggregate validation passed 3/3.

Открытые вопросы

  • 275 "uncertain" speech segments retained without invented wording — automatic transcription is complete, but absolute phonetic verbatimness requires manual listening and must not be claimed yet.
#2 FT09 Local rules/constraints (non-agentic visual logic) — VICTORY, 6 уровней

Durable record of D's second ARC-AGI-3 human trace (FT09): selection rationale, recording, completed evidence-qualified analysis deliverables, the cross-game standardized packaging effort, an operational correction about false completion claims, and a pointer to the third trace (VC33).

Источник: arc-prize-2026-ft09-human-trace.md

Ключевые факты

  • Game slug ft09; live-verified ID 2026-08-24: ft09-0d8bbf25.
  • Archetype: non-agentic visual logic — pattern matching/composition, segmentation/occlusion, relational visual state, click grounding, hypothesis discrimination; no avatar/route/HUD.
  • Video: 1555s (25:55), 3024x1696, ~481 MB.
  • Chronology: 41 continuous intervals (0.000–1554.970s), zero gaps/overlaps; 29/29 evidence refs valid; 3,110 dense frames; 104 overview frames; 249 delta events; ASR span 2.700–1554.010s.

Скиллы — заметки

  • FT09 documents the trace-recording/analysis process rather than gameplay skill sub-rules. Meta-methodology used for extraction: "observation → hypothesis/alternatives → predicted discriminating action → observed delta → support/refute/update", observations kept separate from interpretation.

Факты среды

  • FT09 is characterized as officially non-agentic visual logic involving pattern matching/composition, without an avatar, route, or HUD.
  • Adds segmentation/occlusion, relational visual state, click grounding, and hypothesis discrimination; chosen partly to test whether a DSL/program-synthesis branch is useful.

Разбор по уровням / журнал событий

  • Selected as second human trace after LS20 (LS20 covered navigation/affordances/resources/dynamics/fog/map-memory/phase-planning; FT09 tests non-agentic visual-logic skills).
  • 2026-08-24: D submitted one continuous FT09 screen+voice video, 1555s.
  • 2026-08-25: completion task finished and factually validated. Levels 1-6 and VICTORY!!! visually confirmed.
  • 2026-08-25: normalization task created FT09-event-model packages for LS20 and VC33 plus a common three-game manifest. Aggregate verification passed 3/3.
  • Operational correction (2026-08-25): a prior LS20 report was mistakenly presented as FT09 completion — that claim must never be reused.
  • Subsequently D recorded and completed the third trace, VC33.

Открытые вопросы

  • ASR (speech transcript) was not manually proofread.
  • 2 fps frame sampling cannot prove sub-frame ordering.
  • Level 6 vertical coupling supported by repeated observations but not proven universal.
  • D's required sequence: complete individual traces first, then joint LS20+FT09+VC33 architecture analysis — still pending as of this file.
#3 VC33 Physical/material simulation — Victory pts=794.213 (после доказанного НЕ-случайного gate-клика — старая доктрина "вероятно случайные клики" ОПРОВЕРГНУТА)

VC33 third human trace: two-part 2026-08-24 recording (Part labels reversed vs. content) analyzed into a 429-line evidence-qualified report + CSV timeline, then standardized 2026-08-25, with a decisive gate click at PTS 48.017 driving progress to Victory at PTS 794.213.

Источник: arc-prize-2026-third-human-trace.md

Ключевые факты

  • Part 1 = 873s/14:33 (~317MB); Part 2 = 1000s/16:40 (~390MB); source duration 31:13 total.
  • GAME OVER at PTS 800.013; decisive gate click at PTS 48.017 (lift +50.000, yellow object transfer +52.000); second confirming gate ~PTS 74.000; Victory first frame at PTS 794.213.
  • Standardized package validated: 18 continuous events, 0 gaps/overlaps, 14 evidence paths, 8 decoded JPGs, 38 timeline rows.
  • DOCTRINE CHECK: this file's evidence CONTRADICTS a "probably random clicks, not understood mechanics" characterization — a specific decisive causal click on an identified gate mechanic produced measured, reproducible effects, independently reconfirmed by a second gate.

Скиллы — заметки

  • Transferable agent lessons: trigger a novelty interrupt and affordance sweep for every new visual class; choose tests by budget-aware information gain; separate state predicates from transition operators; represent regions, transfers and gates explicitly as a graph.
  • D became stuck at one point, causing the attempt to run long across two videos — useful reasoning evidence for reconstructing what invalidated/weakened the active hypothesis.
  • The orange conditional gate was noticed but initially misclassified as an indicator; repeated alignment attempts then exhausted the budget, leading to GAME OVER in submitted Part 2 at PTS 800.013.

Факты среды

  • An "orange conditional gate" mechanic: clicking it is decisive — at PTS 48.017 the lift rises by 50.000 and a yellow object transfers by 52.000.
  • A second gate around PTS 74.000 confirms the rule is reproducible, not one-off.
  • The gate was initially misclassified as a mere "indicator" rather than an actionable/conditional trigger, causing wasted alignment attempts and budget exhaustion.

Разбор по уровням / журнал событий

  • 2026-08-24: D submitted two sequential videos; Part labels semantically reversed — submitted "Part 2" contains Levels 1-4 plus GAME OVER, submitted "Part 1" continues Level 4 through to Victory.
  • GAME OVER at PTS 800.013 (submitted Part 2, after repeated failed alignment attempts).
  • Decisive gate click at PTS 48.017 (submitted Part 1).
  • Second confirming gate ~PTS 74.000 (submitted Part 1).
  • First visible Victory frame: submitted Part 1, PTS 794.213.
  • 2026-08-24: 429-line evidence-qualified report and CSV timeline delivered.
  • 2026-08-25: normalized to the FT09 event model; validation passed.
  • Aggregate LS20+FT09+VC33 validation passed 3/3.

Открытые вопросы

— нет данных —

#4 TR87 Program/grammar/sequence — VICTORY, 6/6 уровней, pts=1965.075s

Completed TR87 fourth human ARC-AGI-3 trace (evidence-qualified chronology, reasoning trace, strategy summary, verification), joining the 4-package manifest with LS20/FT09/VC33.

Источник: arc-prize-2026-tr87-human-trace.md

Ключевые факты

  • Video: 1968.679s (32:49), 3024x1674, ~98.1 MiB; single continuous recording to VICTORY!!! (no Part1/Part2 split).
  • 53 reasoning-trace events (TR87-E001–E053), 122/122 evidence references resolved.
  • Per-level attributed spans: L1 478.55s, L2 340.68s, L3 233.40s, L4 350.86s, L5 109.12s, L6 464.70s (longest level).
  • Status counts: 25 supported / 4 refuted / 7 updated / 11 not_tested / 6 not_applicable.
  • Pixel-delta index: 5663 frames; VICTORY spike at pts=1965.075s, mean_diff=16.64 (order of magnitude above any other value — unambiguous).

Скиллы — заметки

  • Board-reading: fixed legend of paired glyphs plus one or two rows of editable glyphs with a white-bracket cursor; confirmed via on-screen HELP text.
  • Navigation/routing: arrows move the cursor along the row and/or cycle the selected glyph's value — consistent across all 6 levels.
  • Goal-setting/prediction: Level 4 required discovering a two-hop indirect "bridge" match (figure matches through an inverted intermediary), discovered only after two full failed scroll-cycles.
  • Prediction/inference: Level 6 (final, ~460s) required recognizing a "1 blue -> 2 yellow via 2 pink intermediaries" structure plus a palindromic 1-2-3-3-2-1 positional symmetry; player revised (1,2,3) to (1,3,2) based on absence of duplicate values.
  • Perception caution: level count HUD briefly misread as "/8" from a small thumbnail, corrected by a tight crop — always re-verify HUD text at full resolution.
  • Perception caution: Level 5's circular grey mask on green was initially misidentified as a distinct "terrain" mechanic, then corrected to the same core game under a different skin.

Факты среды

  • Core mechanic: fixed legend of paired glyphs; below it, editable glyph row(s) with a white-bracket cursor; goal is to match the legend's pairing rule.
  • Levels differ in legend size (6 pairs L1/L2, 8 pairs L3/L4/L6), number of selector rows (1 vs 2), and visual skin (L5 circular grey-on-green vs rectangular grey-on-teal elsewhere — same underlying game).
  • Level 4 introduces a two-hop indirect "bridge" match structure.
  • Level 6 introduces a "1 blue -> 2 yellow via 2 pink intermediaries" structure with palindromic 1-2-3-3-2-1 positional symmetry.

Разбор по уровням / журнал событий

  • Selected as fourth human trace (after LS20, FT09, VC33), per the 2026-08-25 week plan, as the "program/grammar/sequence" archetype-fill trace.
  • Slug tr87 confirmed by browser URL and in-game HUD badge.
  • D submitted one continuous video: 1968.679s (32:49), full playthrough to VICTORY!!! (no split).
  • Full-video pixel-delta index built (5663 frames) to pin exact level-transition PTS and the terminal VICTORY spike.
  • Local faster-whisper ASR run to completion; a separate cloud ASR pass supplied narrative content where needed, spot-checked (~15 key phrases).
  • 6 independent parallel agents (one per level) each read actual frame images and grepped events.csv, producing 53 schema-compliant events merged into one trace.
  • Canonical package written matching FT09/LS20/VC33 standard; manifest updated to 4/4 packages verified.
  • Outcome: full playthrough won (VICTORY!!!) across all 6 levels.

Открытые вопросы

  • Interior event boundaries within levels 3-6 mix exact ASR timecodes with interpolated estimates.
  • The pixel-delta index locates transitions/action clusters but cannot identify which arrow key was pressed.
  • The two-hop bridge rule (L4) and intermediary/palindrome structure (L6) are the player's own successful inferences, not verified against game source code.
  • No separate adversarial "refute this claim" verification round was run for this trace.
#5 TU93 Navigation (глубже, чем ярлык предполагал — predator/vision/chase) — WON 9/9

TU93 human trace (5th completed trace): maze navigation + snake-predator interaction mechanics, WON 9/9 levels, ~23m23s session, confirmed genuine full clear with a stated falsifiable reason at every level transition.

Источник: arc-prize-2026-tu93-human-trace.md

Ключевые факты

  • Session duration: 00:23:23.52 (1403.52s), H.264 3024x1592 @ ~6fps.
  • Archetype: primarily Navigation per pre-play prediction, confirmed but with richer-than-expected mechanic depth (predator vision/facing, chase state, order-dependent elimination, goal-safety rule).
  • 5th completed human trace, after LS20, FT09, VC33, TR87.
  • No doctrine conflicts found; doubles down on doctrine #1 (meaningful pressing) and doctrine #4 (novelty logging).

Скиллы — заметки

  • Local one-step safety fallback under high enemy count (L5, L9): when threat count exceeds what can be mentally tracked, enumerate this turn's legal moves and discard only the ones fatal NEXT turn — cheaper than full simulation.
  • "New indicator that doesn't kill you must do something else" heuristic (L7): when a novel visual change coincides with NOT dying despite apparent danger, promote "enables a different mechanic" over "cosmetic" as default hypothesis, test within 1-2 moves.
  • Novelty-probing method (L4): see new figure → probing moves → infer movement pattern → predict future position → path around it.
  • Order-dependent elimination (L3): three red enemies eaten in specific order (back-to-front, never approaching an un-eaten enemy's eye side).
  • Kiting / indirect obstacle removal (L8): chase-state persists across turns and can be exploited — D lures the chasing enemy in a loop so it commits, opening the original path behind it.
  • Direction/threat encoding via inset sub-pixel (L2): enemy's "eye" side is a directional threat zone active at range 2; approaching from the non-eye side is safe. Learned by trial-death.

Факты среды

  • Grid maze, LEVEL k/9 counter, fixed at 9 levels for this game.
  • Player = blue square ("snake"). Goal = green square. Enemies = colored squares each carrying one smaller inset pixel encoding facing direction = threatened cell(s), not decoration.
  • Magenta/pink horizontal bar fills along the bottom across the run — never narrated by D; likely a per-level/per-run move/time budget indicator, function unconfirmed.
  • Goal-tile safety condition can change between levels of the same game (L5→L6): D stepped onto the goal while a threat was still adjacent and instantly LOST — landing on the goal is only a win if no enemy can eat the player on that same tile/frame, not just present.
  • Orange enemies can move through/overlap red enemies (L6) — red does not block orange.
  • Bordeaux (dark-red) enemy (L7): stationary until player passes near, inset pixel then flips to yellow ("фонарик"), starts chasing every subsequent turn; trigger radius vs facing-cone ambiguous.

Разбор по уровням / журнал событий

  • L1 (0-150s): pure maze traversal, no threats.
  • L2 (~150-300s): first enemy (red) with purple inset eye — died approaching from eye side, restarted, approached from non-eye side, safe.
  • L3 (~300-480s): three red enemies eaten in specific order; first orange enemy appears at level end.
  • L4 (~480-600s): orange enemy movement pattern learned via 2-3 probing moves — straight line, reverses direction.
  • L5 (~600-720s): four orange enemies simultaneously; local one-step safety fallback used.
  • L5→L6 boundary (~660-700s): landed on goal while threat still adjacent, LOST — new "safe on arrival" rule discovered via death.
  • L6 (~720-870s): orange+red together; discovered orange passes through red (one wrong move, then corrected).
  • L7 (~870-1080s): first bordeaux enemy appears, stationary until approached, then chases; inferred via elimination reasoning.
  • L8 (~1080-1200s): same chase-enemy blocks only path to goal; lured via a loop.
  • L9/final (~1200-1400s): combines orange patroller and bordeaux chaser in tighter space; per-turn safety fallback used. Session ends on VICTORY!!!, LEVEL 9/9.

Открытые вопросы

  • Bottom magenta progress bar: never mentioned by D across 23 minutes — move counter, time limit, per-level budget, or cosmetic?
  • L5→L6 goal-tile death: was the earlier level safe-on-arrival by coincidence, or is the rule genuinely different per level?
  • L3 order-of-eating 3 reds: exact decision rule beyond "avoid the one that would eat me" not narrated — trial-and-error or a rule computed in advance?
  • L6 orange passes through red: was there any visual cue before the mistake that could have predicted pass-through behavior, or genuinely unknowable until tested?
  • L7 first bordeaux sighting: exact cell offset at which the pixel flips to yellow — fixed radius-2 trigger or facing-cone check? Ambiguous.
#6 WA30 Sokoban box-placement + coop/adversarial NPC — WON 9/9, "прошёл идеально"

WA30 human trace (2026-08-27): Sokoban-style box-placement game with autonomous helper (orange) and rival (purple) NPCs, WON 9/9 levels, 6th completed human trace, analyzed via the fixed core-skill grid per D's 2026-08-27 doctrine-architecture correction (no new named micro-skills).

Источник: arc-prize-2026-wa30-human-trace.md

Ключевые факты

  • Duration: 1251.33s (~20:51), H.264 3024x1592 @ 4.07fps.
  • HP/move-budget bar levels observed: full at L1/L6, ~90% L7, ~70% L8, ~55% L9.
  • Level-transition timestamps (t≈s): 5, 200, 355(L3), 480(L4), 600(L5), 715(L6), 840(L7), 955(L8), 1080 & 1200(L9), victory 1248.
  • 6th completed human trace, after LS20, FT09, VC33, TR87, TU93.

Скиллы — заметки

  • Perception/board-reading: L1 — proximity to a box changes its render (blur/"spotlight"), the affordance signal for "attachable". SPACEBAR read as context-sensitive via action-probing (not assumption).
  • L3-L5 — reading each helper zone's box-count requirement is a live perception task; D revised a miscount live (open question re: exact process).
  • Movement: L1 — discovers SPACEBAR attach/detach by trial.
  • Prediction of others' moves: L2 — first orange NPC; wrong initial hypothesis (intercept/steal) flipped to "helper, goal-directed" after observing autonomous placement. L6 — first purple NPC; accidentally spacebar-neutralizes it, discovering the "eat a rival" mechanic empirically. L9 — purple stops rather than reroutes when blocked by player's body, blurring the A/B/C NPC classification.
  • Navigation/routing: L3-L5 — player's job becomes pure logistics once walls separate player and helper zones; navigation efficiency named as the only residual risk once mechanic is understood. L8 — new topology confirms L1-L7 rules generalize.
  • Goal-setting/prioritization: L7 — with two purple rivals, D sequences: neutralize both first (pure negative-value blockers), trust orange helpers to self-resolve rest. L9 — delivers to low-risk zone first, then intercepts purple; tracks HP bar to self-estimate "~10 moves left" and reorders actions.

Факты среды

  • HP/move-budget bar — universal across all ARC-AGI-3 games/levels; here visibly depletes level-to-level, confirming it's consumed by moves taken, not just time or deaths (no death occurred this run, bar still drained).
  • Player = green/white square; white side = facing = last movement direction (same "carried directional marker" pattern as TU93, here on the player).
  • Movable pieces = blue-on-black-tile "boxes". Target = blue-gray zone sized to exactly fit the number of boxes required — implicit "how many needed" readout from L1.
  • SPACEBAR is context-sensitive: near a box, attach/detach (carry); near a purple NPC approached from the right side, neutralizes ("eats") it — same input, different effect by adjacency/context.
  • Orange NPC (from L2) = autonomous cooperative ally, independently places boxes.
  • Purple NPC (from L6) = autonomous adversarial rival, carries placed boxes back OUT of the target zone; neutralized by approaching from the correct side + spacebar.
  • Both orange and purple are "type B" figures (goal-directed) but opposite alignment — distinguishing required one observed interaction each, no visual cue (both plain colored squares).
  • L8: the player cannot pass through orange (blocks movement, not decorative).
  • L9: purple respects the player's body as an obstacle — stops advancing rather than rerouting when its corridor path is physically blocked.

Разбор по уровням / журнал событий

  • L1 (0-150s): forms hypothesis "arrange boxes into target zone" without reading Help; regrets not checking Help sooner. Discovers SPACEBAR attach/detach.
  • L2 (~240-360s): first orange NPC; wrong hypothesis corrected after observing autonomous cooperative placement.
  • L3-L5 (~360-720s): multiple orange helpers in walled sub-zones; player becomes pure logistics; D revises a box-count miscount live.
  • L6 (~715-840s): first purple NPC; it removes a placed box; D accidentally discovers the neutralize/"eat" mechanic while trying to reclaim it.
  • L7 (~840-960s): two purple rivals; D neutralizes both first, then lets orange helpers self-resolve rest.
  • L8 (~960-1080s): new map topology, same ruleset confirmed to generalize; new fact logged (can't pass through orange).
  • L9/final (~1080-1248s): two zones, one orange-handled (low priority), one purple-guarded via maze corridor (high priority); delivers to low-risk zone first, then blocks/intercepts purple; tracks HP bar, self-estimates ~10 moves left.
  • Victory screen confirmed on-screen at ~1248s. Result: WON, 9/9 levels, self-described "прошёл идеально".

Открытые вопросы

  • ~780-820s (L6, first purple neutralize): what exactly was the trigger — facing direction, relative position, simple adjacency + spacebar regardless of side?
  • L6-L9, purple's harm potential: was "can purple kill me" ever actually tested/risked, or consistently avoided out of caution?
  • L6-L8, purple persistence: if never eaten, does purple keep undoing placements indefinitely, or stop after some condition?
  • ~480-600s (L3-L4), multi-helper box counts: what's the actual process for reading how many boxes a given orange helper's zone needs, and why did the first read come out wrong?
  • ~1080-1200s (L9), purple stopping when blocked: any visual cue hinting purple might respect an obstacle before testing, or purely trial-and-error? Also: how is the HP/move-budget bar visually tracked to estimate "moves remaining"?

Игра-таксономия

ARC-AGI-3 has 135 environments: 25 public, 55 semi-private, 55 fully private. Of the 25 public games, 3 can be played anonymously (LS20/FT09/VC33), the rest need a free API key. ARC Prize publishes no exhaustive official game taxonomy.

NavigationCoupled kinematics / transportPermutation / construction / reductionLocal rules / constraintsGeometric transformationsPhysical / material simulationProgram / grammar / sequence

Working, unofficial 7-archetype classification from source-code state-transition audit of all 25 public games; may not span OOD mechanics in the private sets.

6. Журнал шагов (хлебные крошки)

  • 2026-08-24 Compute route решён = RunPod (не Kaggle, слабее карта); модель = Qwen3.8-27B (не Qwen3.6, не GLM-5.3-Flash — отклонён D 2026-08-27, офлайн-трек требует ≤1 карты 96GB + без интернета). Первый baseline (v2, честный): LS20 0.0/0-7, ACTION2-деградация 120/130.
  • 2026-08-25 Инцидент idle-billing $10.66 (pod оставлен работать без результата) → D ужесточил дисциплину teardown; позже второй похожий инцидент (~$2.27 за 72 мин почти без работы) → правило: teardown должен быть внешне-supervised, не полагаться только на in-process finally/atexit.
  • 2026-08-25 Архитектура "решено не решать заранее" — сравнить one-brain vs strategist+player vs specialist-role при одинаковых constraints. D поднял вопрос "точно ли 'пройти игру целиком' — правильная веха №1?" → калибровочное исследование → принята лестница: M1 = стабильно проходить уровень 1 (побить внутренний 0.893/1-из-7), M2 = пройти игру целиком (старая цель D, теперь вторая ступень).
  • 2026-08-25 вечер АНТИ-ЛУП ФИКС СРАБОТАЛ: первый уровень когда-либо пройден агентом. LoopGuard (game-agnostic) заменил StuckGuard; LS20 0/7→1/7, score 0.0→0.893.
  • 2026-08-26 Donor-map завершена для 3 self-write пробелов (world-map, object-detector, planner-search) — commodity CS-блоки заимствуются, слой "рассуждение о новых правилах на лету" пишется с нуля, его нет ни в одном репо.
  • 2026-08-26 Round 6 Первый живой прогон 4 игр (LS20/FT09/VC33/TR87), $0, без RunPod: 0/4 уровня. СТРАТЕГИЧЕСКАЯ РАЗВИЛКА перед D: (A) улучшать $0 rules-engine дальше [рекомендация Octo] vs (B) вставить LLM в петлю решений и задействовать RunPod.
  • 2026-08-26 Round 8 D выбрал (B) LLM-в-петле; одобрил ~5-часовой автономный self-repair loop (играть→анализировать→чинить→перезапускать). Round 1: 0/4, но полное время игры подтверждено; найден и зашит action_repeat_stall guard.
  • 2026-08-27 Доктринальная коррекция D (СУПЕРВАЖНО, supersedes все прежние формулировки): отказ от плоского списка микро-скиллов → grid из немногих ядер; HP-бар = универсальный факт среды, не переспрашивать; три TU93 "кандидат-скилла" Octo ОТКЛОНЕНЫ D как не-скиллы. WA30 стала первой трейс-записью, написанной сразу в исправленной рамке.
  • 2026-08-27 D передал решение о СТРУКТУРЕ каркаса нам ("не можешь держать всё в одном окне контекста") → построен STARTHERE + BUILD-LEDGER, ретро-сбор всего диалога начат.

7. Открытые вопросы / пробелы

Доктрина / скиллы

  • Напряжённость про "15-repeat" guard (снять vs единственный proven level-clear механизм) — не решено, флагировать D перед снятием.
  • Purple NPC (WA-30): точный триггер нейтрализации (сторона/смежность?), может ли purple навредить/убить игрока, персистентность если не съеден никогда — все три открыты.
  • TU93: смысл нижнего magenta-бара (сигнал о фильтрах внимания?), точное правило смены goal-not-safe между уровнями, точный decision-rule порядка поедания 3 красных, был ли визуальный намёк на orange-passes-through-red до ошибки, radius-based vs facing-based триггер бордового врага.
  • VC33: доктринальная реплика "вероятно случайные 5-6 кликов, не понятая механика" ПРОТИВОРЕЧИТ содержимому самого трейс-файла (задокументирован конкретный причинный gate-клик с измеримым воспроизводимым эффектом) — флагировать это расхождение D явно.
  • Cognition-bank далеко не у цели ~100-150 Q&A — большая часть работы полки 4 ещё впереди.

Инженерный трек

  • Round 8 LLM-в-петле loop — исход раундов 2+ не зафиксирован в durable на момент последнего обновления каркаса.
  • RunPod pod cjq21kvftu7vaf (round 8) — STOPPED не terminated, решение keep/delete за D.
  • Forgejo/git.octomatica.ru login-recovery тикет d8663003416b — не закрыт.
  • Точная формула scoring соревнования — не зафиксирована из официального SDK.
  • Осталось ~16 из ~22 запланированных видео-трейсов — ждут записи D.