Get in touch
All research

Computer science Technical Report ·

From Evidence to Edits: Rule-Guided Language-Model Video Editing and a Small On-Device Student Model

We gave paper 1’s findings to an AI video editor as rules and tested it on 14 real recordings. With the rules, its edits were about 20 seconds shorter, passed more quality checks and got to the point sooner. A blind AI judge preferred them 7 to 3, but 14 recordings are too few to prove that. We then taught the same rules to Senotel Vision, a tiny model that runs on a phone in about 0.02 seconds with no internet. It follows the rules but doesn’t yet pick the right moments: the judge preferred the large editor on 11 of 13 recordings. The paper reports this negative result openly and explains what it teaches.

Sean Sriphrapradaeng (Senotel)

Read the paper (PDF)

Abstract

A companion report condensed the evidence on what keeps viewers watching short-form video into hook principles, retention principles, genre templates and length bands. Here we ask whether those findings make automated editing better, and whether a small model that runs on a phone can learn to edit as well as a large one. We give the findings to a language-model editor in two ways, as an ordered reasoning brief and as a deterministic post-processor that enforces the mechanical rules on any plan, and evaluate on fourteen openly licensed recordings (52.9 minutes). Compared with the same editor without the evidence, the evidence-guided editor made edits 20.3 seconds shorter on average (p = .010), passed more of six automatic rule checks (p = .014) and brought the first payoff forward from 24.6 to about 15 seconds; a blind, order-balanced language-model judge preferred its edits on 7 recordings and the baseline’s on 3, with 4 ties, a direction too small a sample to confirm (p = .344). We then taught the same material, lesson by lesson, to Senotel Vision-1.0, a 35 KB rule-and-weights editor that plans an edit in about 20 milliseconds without a network connection. It passed the rule checks at least as often as the large editor but did not reach it: it chose the same material far less often than the large editor agrees with itself (F1 0.33 against 0.59), the judge preferred the large editor on 11 of 13 recordings (p = .006), and fitting its weights to the large editor’s choices did not beat the evidence-derived priors on held-out recordings. We report both results in full and describe what the negative one teaches.