YouTube Video A/B Testing: Why MKBHD Is Worried

MKBHD says testing multiple cuts of one YouTube video could fragment comments, timestamps and the shared viewing experience. Here is what the feature does and what remains unclear.

Marques Brownlee speaking on stage at Collision 2023 in Toronto
Marques Brownlee speaking at Collision 2023 in Toronto. Photo by Ramsey Cardy/Collision via Sportsfile, via Wikimedia Commons, licensed under CC BY 2.0. Resized; no other changes.

YouTube video A/B testing worries MKBHD because viewers could receive different edits of the same upload. That could disrupt comments, timestamps and shared discussion. Marques Brownlee is not objecting to testing thumbnails or titles. His concern is that testing the video itself changes the underlying work—and may reward retention without necessarily improving quality.

What did MKBHD say about YouTube video A/B testing?

Brownlee posted his reaction on September 23, 2026, after YouTube announced broader creator tools. In his original statement on X, he said he had thought about video A/B testing since the announcement and liked the idea less the more he considered it. He listed practical questions: Will viewers know they are part of a test? Which version did a commenter see? Will timestamped comments still point to the right moment? How different can the cuts be?

In a longer explanation reported by Dexerto on September 30, Brownlee called the plan an unusually consequential change. His central argument was not that experimentation is always bad. It was that YouTube comments and discussion assume viewers watched the same underlying video.

How would YouTube video A/B testing work?

In its official Made on YouTube announcement, YouTube confirms that creators can compare up to three cuts based on audience attention. Its announcement does not answer every question about comments or version labels.

The proposed feature goes beyond YouTube’s existing title-and-thumbnail tests. Creators would be able to submit multiple edits of a new upload—such as versions with different openings—and compare how audiences respond. Brownlee said creators could review retention data for as many as three versions before one becomes the lasting cut.

According to Brownlee’s account of conversations with YouTube engineers, the platform intends to limit tests to sufficiently similar videos and to the launch period of a new upload. A semantic-analysis system would reportedly reject versions that differ too much. Those safeguards matter, but YouTube had not publicly resolved every viewer-facing detail Brownlee raised when he published his criticism.

The feature sits within a larger push to give creators more optimization tools. The Verge’s September 23 report on Made on YouTube described new automated tools for titles, thumbnails and catalog optimization, including tests designed to identify the strongest-performing version.

Why are comments and timestamps a problem?

Imagine one cut places a product demonstration at 2:10 while another reaches it at 2:45. A comment that says “look at 2:12” may make sense only to people assigned the first version. A reply could appear to contradict the original commenter simply because the two viewers saw different edits.

The problem becomes larger if the winning cut later replaces the test versions. Someone returning to a saved video could find that the sequence, pacing or wording they remember has changed. That is not automatically deceptive—web services routinely test interfaces—but videos function as creative works, references and records in a way that a button color does not.

Does better retention mean a better video?

Not necessarily. Retention measures whether viewers keep watching; it does not directly measure accuracy, originality, usefulness or artistic value. A faster opening may improve a graph while removing context. A more sensational cut may hold attention without making the finished piece more trustworthy.

Brownlee’s warning is therefore also about incentives. If the tool makes retention the visible winner, creators may learn to optimize for the metric even when their real goal is explanation, atmosphere or completeness. That tension is especially relevant for reviewers whose work depends on viewers understanding why a conclusion was reached.

YouTube video A/B testing: rollout and unanswered questions

Several operational details remain unsettled or not publicly documented in full: whether YouTube will label experimental cuts to viewers, how it will preserve timestamp integrity, how comments will map across versions, and exactly how similarity will be judged. Brownlee said he expected a broader rollout in early 2027, but YouTube may refine the system before then.

That makes his critique a warning about design choices, not proof that the final feature will fail. The strongest answer may be transparency: clearly mark tests, preserve version-specific timestamps and let viewers identify which cut they watched.

Why MKBHD’s objection matters

Brownlee has spent years reviewing technology while operating one of YouTube’s most prominent production teams. His concern comes from both sides of the platform: he understands why creators want useful performance data, but he also sees the social cost when optimization changes the object everyone is discussing.

This is a separate issue from his recent public wager over Tesla’s robotaxi timeline, explained in The Daily Vantage’s MKBHD Cybercab bet article. In both cases, his value to viewers is not simply predicting an outcome. It is spelling out the conditions that would make a claim true—and identifying what the available evidence still cannot answer.

YouTube video A/B testing versus title and thumbnail tests

Test What changes? Why it matters
Title or thumbnail The packaging viewers see before watching The underlying video stays the same
Video cuts The opening or edit viewers actually watch Comments and timestamps may refer to different versions

This comparison explains the core of MKBHD’s objection. Better discovery and a different viewing experience raise different questions. His concerns are also discussed in the September 25 Waveform episode.