The test everyone wants to run cannot be run. What a VSL offers instead is a record of exactly where the argument stopped carrying people.

Alex Khassa
Testing a VSL sounds straightforward until you try it properly. Produce version A. Produce version B. Split the traffic. Compare.
That is how testing works for many digital assets. It is not how VSL testing works.
A VSL is an interlocking argument. The opening establishes the problem. The problem creates a reason to keep watching. The explanation creates a reason to believe. The proof supports the explanation. The offer depends on everything before it. Change one section and you may change how every section after it works.
There is a practical problem too. A meaningful variant is not a button-color change. It needs a new script, new recording, new editing, new compliance review and a new communication evaluated on its own terms. And after all that, the downstream event you care about, a qualified inquiry or application or funded relationship, may happen too rarely to make a split test useful in any sensible timeframe.
So the question is not how to A/B test this VSL. It is how to diagnose where the argument is losing people, learn from what happens afterward, and improve the next version.
That takes a different discipline. For how VSLs work generally, see The Ultimate Guide to VSLs for Financial Services.
Usually not, at least not in the clean statistical sense people mean when they say it.
Nothing technically prevents a firm putting two VSLs into two traffic paths. The problem is what happens next.
A useful test needs a variable you can isolate and enough comparable observations to tell whether the difference meant anything. A VSL makes both hard.
Change the opening and you change the context for the problem section, which changes how viewers interpret the explanation, which changes how they receive the proof, which changes how credible the offer feels. By the time someone reaches the call to action you are not measuring one isolated change. You are measuring a different argument. The same applies to replacing the proof, changing the mechanism or restructuring the close.
You can compare two complete VSLs, and that comparison tells you which overall communication performed differently under the conditions it ran in. It does not tell you which sentence, section, claim or visual caused the difference.
Then volume. The meaningful outcome is rarely watched the video. A firm cares about a qualified inquiry, a completed application, a booked appointment, a new account, a policy or an investment relationship, and those sit far down the funnel, affected by audience quality, offer fit, landing page friction, scheduling, follow-up, sales execution and the firm's own qualification process.
Which makes a VSL badly suited to the idea that one version wins because a dashboard shows a better number. The answer is not to stop measuring. It is to measure the part of the experience a VSL can actually reveal.
Where viewers stopped continuing, which lets you diagnose the argument rather than observe a final conversion number.
A VSL gives you something a landing page cannot: a continuous record of the viewer's journey through the argument. A page tells you someone arrived, interacted and completed an action. A VSL shows how the audience behaved as the communication unfolded.
Which creates a diagnostic map. Say the VSL opens on the problem, viewers continue through that section, and the curve changes sharply when the script introduces the mechanism.
That does not prove the mechanism is wrong. It gives you a strong reason to investigate that transition. Maybe the mechanism is confusing. Maybe the explanation asks viewers to accept too much too quickly. Maybe the language turns promotional after an educational opening. Maybe the speaker starts explaining something the audience does not recognize as relevant to them.
The curve does not tell you which explanation is right. It tells you where to look, and that distinction matters because analytics are not mind reading. A drop-off point is evidence of a behavior change rather than a transcript of anyone's thoughts.
It remains unusually useful because the VSL has a known sequence, so you know exactly what argument the viewer met immediately before the behavior changed. You are not running a cleaner experiment. You are getting a more detailed diagnostic record.
The curve becomes far more useful once the script is mapped to the video's actual timeline.
Do not read it as an abstract line. Divide the VSL into functional sections: opening and problem identification, audience qualification, explanation of the problem, introduction of the mechanism, evidence or proof, objections and risk reduction, the offer, and the call to action.
The structure varies by category. A lender explains eligibility and process. An insurer establishes a coverage problem. A fintech explains a product mechanism. An advisory firm explains a planning approach without unsupported promises. The point is not forcing every VSL into one formula. It is knowing what the viewer is being asked to understand at each point.
Then compare that map against the curve. A gradual decline while the argument progresses suggests no single obvious failure. A sharp change at one transition marks a section for investigation. Viewers staying through the explanation and leaving when the commercial ask begins suggests the transition from education to action rather than the education itself.
So review the video and the script together. Put the timeline beside the written argument, mark where the curve changes, then ask what the viewer was being asked to believe at that moment. Considerably more useful than staring at a retention percentage and declaring the VSL good or bad.
A gentle decline suggests ordinary attrition. A cliff suggests a specific moment worth investigating.
Every long communication loses viewers. Someone gets distracted. Someone decides it is not relevant. Someone got the information they wanted. Someone closes the browser. A gradual downward curve is compatible with a VSL that works.
A cliff means behavior changed noticeably around a particular point, which makes that section a candidate for diagnosis.
Look at the transition rather than the sentence at the exact moment of the drop. Sometimes the problem is what was just introduced. Sometimes it is what was removed. A speaker moves from a concrete example into abstract terminology. A script moves from the viewer's situation to the company. A clear explanation suddenly becomes a dense one.
A plateau is informative too, since a curve stabilizing through a section suggests the material holds attention. It does not prove the section persuades.
Which is the important limitation. Retention is not persuasion. Someone can watch because it is interesting and never act. Someone else can leave early because they already understood the proposition and went looking for the offer. So the curve should guide questions rather than settle them.
Because watching tells you the communication held attention, not that the viewer believed the argument, fit the offer or trusted the next step.
This is where VSL testing misleads. Viewers continue deep into the video, the curve looks healthy, and few take the action.
The VSL may still have the problem. The audience understands the argument and does not find the offer relevant. The call to action is vague. The viewer does not know what happens after clicking. The next step feels too large. The offer does not match the problem established at the start.
Or the problem is somewhere else entirely. The traffic is poorly qualified. The page creates friction. Scheduling is inconvenient. The form asks too much. Follow-up is slow. The sales conversation fails to convert appropriate prospects.
Which is why the end of the curve cannot answer the whole question. Analytics tell you viewers reached the call to action and cannot tell you why they did not act on it.
So extend the chain beyond the video. Ask whether the right people watched, what they did afterward, and what happened in the sales process, then connect that back. That prevents the common mistake of rewriting an adequate VSL because a downstream sales problem got assigned to the video.
For the different limitations on the page side, see How to Test Landing Pages for Financial Services.
It is the most practical component to isolate, because it can often be replaced without rebuilding the argument underneath.
The opening has one job: earning enough attention for the viewer to enter the argument. Which makes opening concepts a separate testing problem.
Create alternatives leading into the same underlying VSL. One begins with the audience's problem. One begins with a common misconception. One begins with a situation the audience recognizes. One establishes the consequence of leaving the problem unresolved. The body stays substantially the same.
Not a perfect experiment, since different openings change expectations about everything that follows. It isolates a meaningful component far better than rebuilding the whole argument for every test, and it is more practical to produce, since an opening can be recorded separately and edited into the existing structure.
The key is judging it on whether it brings the right audience into the argument, not on which produces the largest initial spike. An opening attracting curiosity from people unlikely to qualify creates a misleading signal, because the job is establishing relevance rather than entertainment.
Sales conversations reveal what analytics cannot: what prospects believed, misunderstood, questioned or still needed explained.
The most valuable source of VSL feedback, and the most underused.
Talk to the people who speak with prospects. What do they say immediately after watching? What do they think the company does? What problem do they believe is being solved? What do they ask to have clarified? Which objection appears repeatedly? What expectation did the VSL create that the sales team now has to correct?
Those conversations expose what a curve cannot. A viewer can watch almost the entire VSL and still misunderstand one important point, and the curve will never show it. A prospect can arrive believing the service includes something it does not, which is an expectation problem rather than a retention problem. A prospect asking repeatedly about the same missing piece of information is telling you the VSL needs a clearer explanation.
Sales calls also reveal when the VSL is working. If prospects arrive already understanding the problem, the approach and the next step, the team spends less time explaining the proposition and more determining fit. A meaningful outcome even though it never appears as a testing metric.
Which matters more in this category, where the decision involves trust, suitability, risk, eligibility and a significant financial commitment. The sales team is hearing the questions analytics cannot hear.
Replacing one version with another over time is usually more practical than splitting traffic between expensive productions.
Call it sequential replacement rather than A/B testing, which is the honest name for it.
Run the current VSL. Gather the curve. Review the sales conversations. Identify the strongest suspected weakness. Produce a revised version addressing it. Put that into the programme and compare what happens over the following period.
The advantage is that it reflects how VSL production actually works: learning from one complete argument and using that to build the next.
The limitation is equally important. The versions are not running under identical conditions at the same time, so you cannot attribute every difference to the VSL. Traffic changes. The market changes. The offer changes. Seasonality changes. Sales execution changes. The audience becomes more or less familiar with the firm.
So sequential replacement produces directional learning rather than laboratory certainty. Which sounds less impressive than a clean A/B test and is considerably more honest. The goal is not manufacturing certainty from weak data. It is making better decisions with the evidence the programme can realistically produce.
The part where the strongest evidence points to a specific communication problem, starting with the earliest meaningful failure.
Do not rewrite the whole VSL because one section looks weak. Start with the evidence.
Viewers leaving sharply when a concept is introduced means examining whether it is confusing, too abstract or insufficiently connected to the problem. Leaving during a transition from education to company material means examining whether the communication turns promotional. Staying through the explanation and abandoning around the call to action means examining the transition, the offer clarity and the expectations around the next step.
A reasonable curve alongside sales conversations revealing repeated confusion means addressing that specific misunderstanding in the script. A VSL holding attention while the traffic is poorly qualified means investigating the acquisition source before touching the video. And an objection prospects raise repeatedly that the VSL never addresses means deciding whether it belongs in the argument.
Fix the biggest supported problem first, and do not change five things at once then pretend you know which mattered. Then watch what happens in the next version, which creates a learning loop without pretending every iteration is a controlled experiment.
The goal is not proving one VSL is universally better. It is building enough evidence that the next one communicates more clearly to the right audience.
Digital marketing wants every optimization problem reduced to a winner and a loser. VSLs resist that treatment, because performance depends on the audience, the offer, the speaker, the argument, the traffic source, the surrounding funnel and what happens after the video.
So the evidence comes from several places. The curve shows where attention changes. The argument map shows what the viewer met at that point. The sales team shows what qualified prospects understood or misunderstood. Downstream behavior shows whether attention became action. The next version shows whether the change improved the diagnosed problem.
None of those is perfect on its own. Together they build a stronger decision process than a forced A/B test producing a result that looks precise and answers the wrong question.
Which matters here because the objective is rarely maximum watch time. It is a communication that attracts an appropriate audience, explains the proposition clearly, sets accurate expectations and moves the right people toward the right next step.
Every meaningful variant is a new communication that goes through the firm's own review before it is used.
Testing does not create an exception to compliance obligations.
Changing a headline, claim, example, proof point, testimonial, disclosure, product description or call to action can create a different communication requiring review under the firm's procedures. For firms subject to the SEC Marketing Rule, applicable requirements depend on the nature of the communication and the claims, endorsements, testimonials or performance information involved, with the firm's own compliance professionals and counsel determining what applies.
Which has a practical consequence. Compliance is not something that happens once at the start of a VSL programme, because iteration creates new review requirements.
So build it into production. Document what changed between versions. Keep a clear review history on the script. Retain supporting evidence. And do not casually add claims because a testing hypothesis suggests they might improve response.
Most of all, do not assume a variation is acceptable because a previous version was approved. How to Make VSLs That Pass Compliance covers that in detail, and the principle is the same: no outside article determines what your compliance team approves.
A VSL is not video production with analytics attached. It is an argument delivered through video, which means useful testing starts with the argument.
What problem is being established, and for whom? What does the viewer need to understand before the solution makes sense? What evidence supports the explanation? What objections need addressing? What do they need to believe before taking the next step?
Then interpret the analytics against those questions. Leaving at a section means investigating that section. Staying but arriving on sales calls confused means investigating the explanation. Understanding the proposition and not acting means investigating the offer, the next step, audience fit and the surrounding funnel. An opening attracting attention from the wrong people means optimizing for relevance rather than engagement.
And when a new version looks stronger, resist declaring a scientific winner when the evidence does not support one. That restraint is not weakness. It is what makes the process credible.
A VSL does not give you the clean A/B test every marketer wants. It gives you something different: a record of where the argument stopped carrying the audience forward. Used properly, that tells you where to investigate, what to ask the sales team, and what to change next time.
Enough to build a disciplined programme, and a far better foundation than manufacturing statistical confidence from a test that was never capable of answering the question. For constructed examples of VSL structures, see VSL Examples for Financial Services.
Book a call and we'll walk through the math for your firm. How many appointments you'd need, what the unit economics look like, and whether we're a fit.
Install the AUM OS in your firm today and scale up with virtual appointments.
Answers based on what we've seen drive top performance across years of data.
First appointments typically hit the calendar within the first 1–2 weeks after launch. Month one is optimization. Month two is when things stabilize and become predictable.
2–3 hours of video recording every 3–6 months. That’s it. We handle everything else.
We’ve worked with over 200 RIAs and their compliance departments. We know what gets approved under Special Ad Category restrictions. We build compliant from the start and coordinate directly with your team.
Total marketing budget starts at $17,500 per month and ranges up to $120,000 depending on your goals, ad spend included. Engagements run on a 12 month minimum.
No. And you should be skeptical of any agency that does. Guarantees in this space are a red flag — they’re selling you a feeling, not a strategy. What we offer is a proven methodology, a team that’s managed over $10 million in Meta ad spend for RIAs, and a track record of $45+ Billion of AUM pipeline generated across 200+ firms. The firms that follow our methodology and commit to the process see results. That’s why we’re selective about who we work with.
Most agencies try to do everything — Google, email, social, websites — and they’re mediocre at all of it. We only do Meta Ads for financial firms. We’ve spent over $10 million in this exact channel under Special Ad Category restrictions. We know what works because it’s all we do.
Good. Most of our clients do. We’re not replacing your marketing person or your agency. We’re adding the one capability they probably don’t have: Meta Ads at scale with branded video for financial services under Special Ad Category. We plug in alongside whatever else you’re running.
No. We do Meta Ads. That’s our entire focus. If you need those other services, we’re happy to recommend partners, but that’s not what we do.
Click the button below to apply. If it’s a fit, we’ll schedule a strategy session to walkthrough timelines, pricing, and how AUM OS would work for your firm.
Install the AUM OS in your firm today and scale up with virtual appointments.