Video doesn't go in
Claude and GPT accept no video at all. Tools cut it into frames, usually one per second.
Video perception for language models · MCP
Claude and GPT don't take video. Frame samplers drop the flash, the cut, the glimpsed object, and the sound. Framegaze lets a language model actually watch.
01 — The problem
Claude and GPT accept no video at all. Tools cut it into frames, usually one per second.
The audio usually disappears entirely: no speech, no laughter, no bang.
Even models that hear audio look at only 1–2 frames per second. A flash, a 2-frame insert, or an object crossing between samples simply doesn't exist.
02 — How it works
Each frame gets a cheap check. Expect the next frame to be the last one shifted by motion; whatever motion can't explain is surprise. It is measured per cell and normalized to that cell's habit, so a speck in a corner stands out even on busy footage.
Speech with words and timestamps (Whisper). What is making noise — laughter, applause, sirens, music (CLAP, zero-shot). Loudness. And attention per frequency band: a steady beat doesn't surprise, a new sound does.
The model receives a protocol of what is seen, heard and said, and when; one storyboard image of the whole clip; and contact sheets before, at, and after each event.
When something needs a closer look, the model asks for it — dense frames of a span, the sound and words of a stretch, or a single frame up close.
03 — Benchmark
04 — Works where you work
Runs on CPU on your own machine. Nothing is uploaded. Ships as an MCP server for Claude Code, Claude Desktop and any MCP client.
05 — Where it helps
Find stray flashes, flash frames and glitches before the cut ships.
Catch the toast that appears for a moment and the error dialog that blinks past.
Analyze fast cuts, on-screen text and sound cues the way a viewer meets them.
Review what was said, what was shared on screen, and when.
Early builds go to a small group first.
Your mail app should open with a ready message. If it doesn't, write to [email protected].