A lot of my work is watching surveillance footage and writing down, second by second, who does what: which car arrives, who gets out, which door opens, who leaves first. For the last few months the first draft of that timeline has come from Twelve Labs Pegasus 1.5, a model built specifically for video. Then, within five weeks, Google released Gemini 3.8 Flash (2 September) and Twelve Labs released Pegasus 1.6 (6 October). There is no public benchmark that covers all three, so I ran them on my own footage and scored them against timelines I had verified frame by frame.

The test set

34 minutes of real DVR and IP-camera recordings from my archive, chosen to be the kind of footage that makes video models fail: night, infrared, motion-triggered recording, people far from the camera.

Camera Conditions Key moments What happens
Car wash evening, colour, close range, 1080p at 15 fps 10 a car arrives and parks, the driver gets out and leaves, comes back, leans in through the front passenger door, closes it, gets in and reverses out
Yard at night infrared, then colour, mid-distance, two gaps in the recording 19 two cars, two men, doors opening and closing, lights switching on and off, manoeuvres and departures
Forecourt at night colour, motion-triggered recording, 12 fps 4 a car crossing the frame twice, a person standing by it in the far background
Street at night colour, people about 90 pixels tall at the far end of the frame 13 a scuffle: a push from behind, a fall, a kick attempt, a swing, two groups separating; a police car
Car wash, 10 minutes as one clip as above 10 the same moments, to see how each model copes with long input

That is 46 key moments in the four 2-minute-segment runs, plus the 10 moments of the long-input run. The ground truth is my own timeline, checked on the individual frames.

The contenders

Model Settings Price
Twelve Labs Pegasus 1.5 pegasus1.5, temperature 0.2 $1.75 per hour of video + $7.50 per 1M output tokens
Twelve Labs Pegasus 1.6 pegasus1.6, temperature 0.2 same as 1.5
Gemini 3.8 Flash, default 1 frame per second, default resolution (about 70 tokens per frame), thinking medium $0.75 per 1M input tokens, $3.75 per 1M output tokens (including thinking)
Gemini 3.8 Flash, high resolution 1 frame per second, MEDIA_RESOLUTION_HIGH (about 270 tokens per frame), thinking high same
Gemini 3.8 Flash, high resolution, 3 fps 3 frames per second, high resolution, thinking high same

Gemini's prices are introductory: on 1 January 2027 both input and output double, to $1.50 and $7.50 per million tokens. The whole round cost about $5.

How I scored

Every model got the same prompt and byte-identical clips. The clips are 2-minute stream-copied segments, the format my pipeline already uses. The prompt asks for times as elapsed time within the clip, never read from the burned-in clock, one event per line. It also asks the model to keep what it sees separate from what would merely be consistent with it: two overlapping silhouettes are not a blow. Each key moment then got one of three marks:

  • Correct: reported within ±6 seconds, with the essential detail right.
  • Partly: the moment is there, but a key detail is wrong or missing. Typical cases are the wrong door, "walks in from the left" instead of "gets out of the car", or "interacts with the car" instead of what he actually did.
  • Missed: not reported, more than 15 seconds off, or contradicted.

Separately I counted material errors: false statements inside the verified windows, such as a car described as parked after it had driven away. I checked every disputed claim on the frames. One caveat: I was the only judge and I knew which model wrote what. Forty-six moments from four cameras is a small sample, and a difference of a few points is noise. The qualitative differences below are the part I would bet on.

Results

The four cameras in 2-minute segments. Score = (correct + ½ × partly) / 46.

Model Correct Partly Missed Score Material errors $ per video minute (now) $ per video minute (2027) Seconds per 2-minute segment
Pegasus 1.5 17 8 21 46% 7 0.031 0.031 95
Pegasus 1.6 16 10 20 46% 6 0.031 0.031 20
Gemini 3.8 Flash, default 20 8 18 52% 4 0.007 0.013 24
Gemini 3.8 Flash, high resolution 23 6 17 57% 3 0.027 0.055 66
Gemini 3.8 Flash, high resolution, 3 fps 23 10 13 61% 4 0.055 0.110 94

Per camera (correct / partly / missed):

Model Car wash (10) Yard at night (19) Forecourt (4) Street scuffle (13) 10-minute clip (10)
Pegasus 1.5 4 / 3 / 3 9 / 5 / 5 3 / 0 / 1 1 / 0 / 12 2 / 1 / 7
Pegasus 1.6 5 / 3 / 2 7 / 7 / 5 3 / 0 / 1 1 / 0 / 12 2 / 3 / 5
Gemini 3.8 Flash, default 5 / 3 / 2 11 / 5 / 3 3 / 0 / 1 1 / 0 / 12 3 / 3 / 4
Gemini 3.8 Flash, high resolution 5 / 3 / 2 13 / 3 / 3 4 / 0 / 0 1 / 0 / 12 6 / 1 / 3
Gemini 3.8 Flash, high resolution, 3 fps 5 / 3 / 2 13 / 4 / 2 4 / 0 / 0 1 / 3 / 9 5 / 2 / 3

On the close, well-lit car-wash camera all five are practically equal. The differences come from the night yard, from small details far from the camera, and from long input.

What each model gets wrong

  • Pegasus 1.5 falls back to filler when the footage gets hard. Through the exact window of the street scuffle it produced ten-second buckets of "A car drives past the camera" and "A person is walking on the sidewalk". On the night yard it wrote seven identical "a dark sedan crosses the foreground" lines in two minutes, when one car did. It turned a man leaning in through the front passenger door into "opens the rear hatch and places items inside". It is also the slowest, at about 95 seconds per 2-minute segment.
  • Pegasus 1.6 is exactly as accurate as 1.5 and five times faster. It is noticeably better at describing close, well-lit people: two passing pedestrians got their clothing item by item and the bag in the left hand, all correct. It also says "rear hatch" for the passenger door, and it declared a car "remains parked with its lights off" after that car had driven away. On the 10-minute clip it turned a blue sedan into a pickup truck and had the attendant climb onto its bed. The new features in 1.6 target first-person (robotics) video, which this test does not cover.
  • Gemini 3.8 Flash at default resolution is already ahead of both Pegasus versions on the night yard and costs a fifth as much per minute. It misses small and distant things.
  • Gemini 3.8 Flash at high resolution noticed a person standing next to a parked car at the far end of the forecourt. Only the two high-resolution runs did; I had not written it down myself and confirmed it on the frames afterwards. It caught the driver getting out of the car and back in, and a jump in the recording ("An abrupt cut occurs in the video footage") that no Pegasus run mentioned. Its weak spot is sides of the car: twice it said "driver's door" for a passenger-side door, and once "moves forward" for a car that reversed out.
  • Gemini 3.8 Flash at 3 fps scores best and was the only run that saw a parked car's lights go off. It costs twice as much as high resolution at 1 fps and is as slow as Pegasus 1.5. It was also the only run that flagged the scuffle at all: two people running across the road towards a group, with the note that "whether any physical contact occurs cannot be determined".

Nobody saw the scuffle. A push from behind, a fall and a kick attempt, involving people 90 pixels tall at night, were invisible to all five configurations. For that kind of footage the model is a map of where to look, and the frames themselves remain the evidence.

Long input

Pegasus degrades badly on long clips, which is why my pipeline cuts video into 2-minute segments in the first place. On one 10-minute clip of the car-wash camera Pegasus found 2 of the 10 moments, with either version. Gemini at high resolution found 6, one more than in 2-minute segments. At about 16,000 tokens per minute, 10 minutes of high-resolution video is 160,000 tokens, comfortably inside the 1M context. Segmenting is still useful for detail and for parallelism, but with Gemini it is no longer a correctness requirement.

Two lessons that are not about the models

  • Check your clip timestamps before you blame the model. DVR exports in .avi often carry no packet timestamps. A fast stream-copy cut starts at the previous keyframe, which on one camera was up to 13.3 seconds earlier than requested. If your code assumes the clip starts where you asked, every reported time is late by up to that much. My first scoring pass made Pegasus look up to 12 seconds late on that camera. The fix was to read the keyframe positions from decoded frames instead of packet flags.
  • Read the data terms. On Gemini's free tier your uploads are used to improve Google's products and human reviewers may read them. The paid tier does not train on your data and keeps logs for 55 days for abuse monitoring (terms). New Google accounts are on a prepaid plan: linking a billing account is not enough, and the API answers HTTP 402 until you buy credits (minimum $5, valid 12 months). Twelve Labs' self-serve terms let it use your uploads to train its models unless you opt out in the account settings. The same terms forbid uploading content with information about identifiable individuals (terms of use, sections 4(a) and 13(c)(ix)). For surveillance footage, that is worth knowing before you upload anything.

What the public benchmarks say

Very little. The only video benchmark in the Gemini 3.8 Flash evaluation is LVBench (87.8%, up from 85.4% for 3.7 Flash), measured by Google itself. Pegasus 1.5 was launched with Twelve Labs' own segmentation and prompting metrics against Gemini 3.1 Pro, and Pegasus 1.6 with no numbers at all. None of the three appears on an independent leaderboard. If you work with video, your own footage is the only benchmark that counts.

Verdict

For describing surveillance footage, in October 2026, Gemini 3.8 Flash at high media resolution and 1 frame per second is my new default. It is clearly ahead of Pegasus on night footage and small details, does not degrade on long clips, and costs about the same as Pegasus until the end of the year (twice as much from January, which is still about $6.50 for a 2-hour recording). I switch to 3 fps only for short windows where a one-second action matters, and default resolution makes a cheap first pass over hours of footage to find where something happens. If you stay with Twelve Labs, move to Pegasus 1.6: same quality, same price, five times faster.

None of them is a witness. Every model got doors wrong, and every model missed the scuffle. The draft tells me where to look; the frames tell me what happened. The benchmark is scripted, so when the next model appears the rerun is one command and about $5. I will update this post when that happens.