A vision transformer with native 2D spatial awareness for the Abstraction and Reasoning Corpus.
pip install torch numpy tqdm matplotlibcd arc2d
python run_tests.pyWhat to expect: All 8 modules should show ✓ PASS. Time: Under 2 minutes, no GPU needed.
python train_single_task.pyWhat to expect: Loss decreases from ~2.4 to near 0. Accuracy > 90%. Time: 1-5 minutes on CPU.
PYTHONPATH=. python -m visualization.bias_tablesWhat to expect: Heatmap images saved in bias_table_plots/ showing spatial patterns.
Time: Under 1 minute.
git clone https://github.com/fchollet/ARC-AGI.gitPYTHONPATH=. python training/pretrain.py --data_dir ARC-AGI/data/training --size smallOr use synthetic tasks for a quick test without ARC data:
PYTHONPATH=. python training/pretrain.py --synthetic --epochs 20Time: 2-4 hours on GPU (Small config).
PYTHONPATH=. python evaluation/evaluate.py \
--model_path checkpoints/best_model_small.pt \
--eval_dir ARC-AGI/data/evaluation \
--ttt_steps 100 \
--n_views 10Time: ~30 seconds per task.
arc2d/
├── config.py # Model configurations (Small/Medium/Large)
├── run_tests.py # Checkpoint 1: run all self-tests
├── train_single_task.py # Checkpoint 2: single task overfitting
├── requirements.txt # Python dependencies
├── README.md # This file
│
├── model/ # The neural network
│ ├── attention.py # ★ Full 2D Attention (core innovation)
│ ├── ffn.py # ConvGLU feed-forward network
│ ├── block.py # Transformer block (attention + FFN)
│ └── model.py # Complete model
│
├── data/ # Data loading and processing
│ ├── canvas.py # Canvas placement utilities
│ ├── augmentations.py # Geometric + color augmentations
│ └── dataset.py # ARC JSON file loader
│
├── training/ # Training scripts
│ └── pretrain.py # Stage 1: offline pretraining
│
├── evaluation/ # Evaluation scripts
│ └── evaluate.py # TTT + multi-view inference
│
└── visualization/ # Visualization tools
└── bias_tables.py # Checkpoint 3: visualize bias tables
Input: ARC grid (integers 0-9) placed on a fixed-size canvas
Three components:
- 2D Attention — per-head learnable bias tables indexed by (Δrow, Δcol)
- ConvGLU — depth-wise 3×3 convolution on the gate path
- Output head — per-pixel classification into 11 color classes
No positional embeddings. Position enters exclusively through the 2D bias tables.
Training: Pretrain on 400 ARC tasks + augmentation, then test-time training per task.