Repository navigation
Gemma 4 hybrid continuous-state reasoning #795
Larkooo
started this conversation in
Show and tell
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Hello everyone! I've been experimenting with a Coconut-inspired continuous-input path for Gemma 4 E2B.
The experiment replaces an initial text reasoning step with two learned continuous feedback
positions, then generates the remaining reasoning and final answer. The loop uses
no token sampling or nearest vocabulary lookup, preserves Gemma 4's continuous
per-layer embedding branch, and zeros the token-indexed contribution at latent
positions!
This experiment builds on Coconut (Hao et al.). I adapted its continuous-feedback approach for Gemma 4 E2B, using a learned feedback bridge and Gemma-specific embedding handling.
I used two latent steps as a starting point. The count is configurable, but this
checkpoint was trained with two. I also have experimented with one additional latent step with inference taking 5.3% longer by median latency but no apparent gain in accuracy.
Here are some benchmarks I ran on 600 held-out diagnostic arithmetic and
link-traversal test questions:
In this run, the hybrid answered 583 out of 600 questions correctly and reduced median
completed-answer latency by 10.6%. Mean latency fell by 16.8%. That's a
2.5 percentage-point accuracy loss for a faster final answer.
I also evaluated the same checkpoint on 400 harder questions. The hybrid scored 360/400 (90.0%) versus 389/400 (97.25%) for text reasoning. On larger arithmetic operands, accuracy was 80.5% versus 95%; both methods reached 99.5% on longer link chains. The hybrid hit the output limit on 11 questions. The small accuracy loss from the original test therefore didn't hold on harder arithmetic. Results and raw measurements.
I haven't established whether the feedback content provides an advantage over simpler extra computation positions. On 100 validation questions, normal feedback scored 96/100 and zeroed feedback scored 93/100, with a paired difference interval spanning zero. A separately trained zero-position control scored 66/100, but that was one training seed and its validation loss was still improving when training stopped.
Generating less reasoning text reduces sequential decoding work, but the latent positions and transition also cost computation. The measurements show a latency tradeoff on these tasks; they don't establish lower total FLOPs or a general improvement in reasoning.
I measured these runs on my MacBook Pro M5 with 32 GB of memory. These timings describe the recorded sessions, and I still need a controlled repeat under idle-machine conditions.
I've shared the code, measurements, and limitations. I'd love feedback on the controls and whether the Gemma-specific continuous-input support would be useful as a community example. If there's interest, I'd be happy to lead that contribution!
All reactions