We ran our own tool on a video about being found by AI
A YouTube video explaining how AI search decides what to cite is mostly a person talking over slides. The slides carry the examples. We pointed aitubenotes at one and kept what came back, because the gap between what was said and what was shown is the whole argument for reading both.
What is the difference between SEO, AEO and GEO?
SEO makes a site discoverable through on-site, technical and off-site work. AEO structures content so an answer engine can extract a direct answer, such as a featured snippet. GEO builds structure, content and authority so generative systems see, cite and recommend you. They stack rather than compete: the video puts 76% of AI overview citations in the top ten conventional results, so the base layer still decides the rest.
What did reading the screen catch that the transcript missed?
The actual words of the worked example. Explaining how not to open a page, the speaker waves the sentence away as "yada yada yada" and moves on. The full sentence is on the slide behind him for the whole passage. A transcript-only tool records the dismissal. The notes record the example.
And a bad example of this is answering that right away, saying, "In today's digital landscape, businesses are increasingly looking for ways to improve yada yada yada."
Let's say youre writing a "In todays digital landscape; businesses are increasingly looking for ways to improve their online presence SEO has evolved significantly over the years, and in this comprehensive guide explore.
That is reproduced exactly as the pipeline produced it, missing apostrophe and all. Optical character recognition on a compressed video frame is not perfect, and we would rather show what it really returns than a cleaned-up version. The example survives the noise, which is the part that matters.
Why does a slide carry more than the words spoken over it?
Because a presenter narrates what is on screen instead of reading it out. Diagrams, dashboards, code and comparison tables get shown and then referred to as "this" or "here". Every one of those pronouns is a hole in a transcript. Introducing a worked site structure, the speaker calls it "a bit of an architecture"; the diagram itself is on the screen.
let's say for this instance, you're a local business. You're a plumber in Austin, Texas.
Website Structure: Austin Plumber Ideal architecture for Local SEO + GEO optimization Total indexed pages: 137 SEO Foundation AEO Content (Capsule Technique) GEO Optimization (Al Citations) Homepage MustIMPLUVbeaprO.
The scale the example operates at, 137 indexed pages, appears nowhere in the audio. Neither does the layered shape of the diagram. Anyone working from a transcript would come away knowing that a plumber was mentioned.
How should a page be written so an AI system can cite it?
Put a question in the heading and answer it in the first sentence underneath, then elaborate below. The test the video gives is whether that sentence still makes sense lifted out on its own. The named failure is opening with scene-setting instead: its example of what not to write begins "In today's digital landscape", and assistants skip past it.
Why should distinct services get distinct pages?
Because one page cannot answer two search intents, and combining them makes the services compete with each other. The guidance is to give each service its own page and differentiate the content by title tag, meta description, H1, H2s, opening sentence and schema. Schema is singled out as "the one you cannot miss out", and it should describe the specific service on that page.
Does this only apply to marketing videos?
No, and marketing videos are the easy case. Conference talks, recorded lectures, code walkthroughs and product demos are all mostly visual: the terminal, the whiteboard, the chart and the diff are the content, and the speaker talks around them. Anything you would pause and screenshot is a video whose transcript alone loses the substance.
What we will not claim
Optical character recognition on video frames is imperfect, as the quotes above show. Handwriting, dense small text and low-resolution uploads read poorly. Nothing here is reviewed by a person before it reaches you, so the notes carry a line saying a machine wrote them, and anything that matters should be checked against the source. This is one run on one video, and it is not a benchmark.