---
title: "Gemma 4 12B deleted the encoders"
date: 2026-06-03
series: 10
summary: "Gemma 4 12B just deleted the vision and audio encoders. Most multimodal models still bolt on three separate networks. Google ships one weight space."
voice: builder
tags: [encoder-free-multimodal, unified-embedding-projection, native-audio, on-device-16gb, gemma4]
image: "/notes/gemma4-encoder-free/image.png"
imageAlt: "split panel: legacy 3-encoder stack vs Gemma 4's single unified weight space"
linkedin: null
sources:
  - title: "Encoder-free architecture, 48×48 patch projection, audio linear-projection, dropped conformer layers — MarkTechPost"
    url: https://www.marktechpost.com/2026/06/03/google-deepmind-releases-gemma-4-12b-an-encoder-free-multimodal-model-with-native-audio-that-runs-on-a-16-gb-laptop/
    date: 2026-06-03
  - title: "Apache 2.0, runs on 16GB laptop, native text/image/audio/video — VentureBeat"
    url: https://venturebeat.com/technology/googles-new-open-source-gemma-4-12b-analyzes-audio-video-and-runs-entirely-locally-on-a-typical-16gb-enterprise-laptop
    date: 2026-06-03
  - title: "Official launch (Google did not publish full benchmarks at launch) — Google"
    url: https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/
    date: 2026-06-03
dateApprox: false
---

Gemma 4 12B just deleted the vision and audio encoders. Most multimodal models still bolt on three separate networks. Google ships one weight space.

How it actually works:

**Vision.** Raw images → 48×48 patches → a single matrix multiply into the LLM's embedding space. No standalone ViT, no cross-attention adapter.

**Audio.** Raw 16kHz frames → linear projection into that same text embedding space. RoPE already handles the timing, so the 12 conformer layers older Gemma needed are simply gone.

**Result.** Three encoder stacks collapse to one. Apache 2.0, Q4, runs on a 16GB laptop. Text, image, audio and video in native.

Google published no benchmarks at launch. The "near-26B at half the memory" line going around is community-reported — treat it as a hypothesis, then benchmark your own task.

**The trap:** one weight space means your single Q4 build *is* the whole multimodal stack, and LoRA now updates vision, audio and text in one pass. Great — until you quantize. There's no separate encoder to keep at higher precision, so a sloppy quant degrades every modality at once.

When the encoders disappear, your quantization config becomes your entire multimodal strategy.
