The same forty-two minute talk is four different documents depending on why you sat through it.
VideoSage turns a video into something you can use. The obvious version of that product is a summariser: paste a link, get paragraphs. Building toward it exposed the assumption underneath, which is that summary is a format. It is not. It is a word covering at least four different jobs that have almost nothing in common.
A strategist watching a conference talk wants the argument and the evidence. A team lead wants the practices they could adopt on Monday. Someone studying wants to be tested on it. Someone running a content account wants three posts out of it. One video, four intentions, and a single generic summary serves none of them well.
Summary Hides the Question: Every summary answers an unstated question. Leaving it unstated means the system picks one on the viewer's behalf and never says which.
The Verification Gap: A summary you cannot check against the source is a rumour with good formatting. The source is an hour long, so checking has to cost seconds or nobody does it.
Filtering Is Invisible by Default: Condensing is mostly deciding what to drop. Those decisions shape the output more than anything the model writes, and they usually happen where the user cannot see them.
The format choice comes before the processing, not after
The first screen is not a text field or a progress bar. It asks what the output should be: a summarised report with key takeaways and timestamps, a business playbook of applicable strategies, a learning module with quizzes and topic deep-dives, or content automation that produces threads, blog drafts and notes.
Putting that choice first changes what the system is doing. It is not producing a summary and then styling it four ways. The intent determines what gets extracted in the first place, because a playbook needs actions and a learning module needs testable claims, and those are different passes over the same transcript.
It also makes the interface honest about its own behaviour. The user has stated the question, so the output can be judged against it rather than against some general notion of accuracy.
Timestamps are the part that makes it checkable
Every section of the generated report carries the time range it came from. An introduction covers 00:00 to 05:32; the section on key technologies picks up at 05:33.
That is the smallest feature here and the one doing the most work. It converts each claim from something the system asserts into a pointer at the source, which means disagreement is cheap: jump to 12:04 and hear what was actually said. Without it a viewer has two options, believe the whole thing or rewatch the whole thing, and one of those defeats the product.
It is the same principle I keep arriving at in AI work from different directions. The interface does not need to prove the model was right. It needs to make checking fast enough that being wrong is survivable.
Making the filter visible
Alongside the output sits a panel of focus areas and highlight themes: key takeaways, actionable insights, deeper analysis, concept explanation, and a set of subject themes from leadership to finance.
These are filters, and putting them on screen next to the result is a deliberate choice to expose what would normally be a hidden editorial decision. The output is a view of the video under a stated set of priorities rather than the definitive account of it. Regenerate is a primary button for the same reason: a summary is a view you can change your mind about, not a document the system hands down.
The design risk is obvious. Given controls, people will tune a summary toward what they hoped the video said. Timestamps are the counterweight, since a filtered claim still points back at the moment it came from, but I would want to watch someone use both together before claiming the balance is right.
Where it stands
VideoSage is the least finished thing on this site. The interface exists as a design, no model runs behind it, and the entry for it in my projects list was an empty shell for months because I never wrote it up.
I am keeping it because the format-first idea is the one I would take into a real product. Ask what the output is for before deciding what to extract, and point every claim back at the second it came from. Whether four formats are the right four, and whether people actually choose before they see anything, are both open.
