September has always been a month of restarts. For digital teams, however, restarting has become more complicated than reopening Jira and picking up the first item in the backlog.
Release cycles are shorter, products evolve continuously, customer expectations shift and digital experiences increasingly depend on complex ecosystems of devices, operating systems, integrations and third-party services. AI is adding another layer of acceleration, making it possible to build, test and iterate faster than before.
Against this backdrop, a roadmap agreed before summer can still be perfectly valid in September. But it should not be treated as automatically valid.
A backlog records what a team decided to do. It does not necessarily tell us whether the evidence behind that decision is still current, whether the problem has changed or whether a previously reasonable assumption has become a source of risk.
So before asking “What should we work on first?”, there is another question worth asking:
What would have to be true for our Q4 priorities to succeed and how certain are we that it still is?
Every product decision contains assumptions.
A team redesigning onboarding may assume that reducing friction will increase completion. A team working on checkout may believe a particular step is responsible for abandonment. A Product Manager prioritising a new feature may be relying on recurring customer requests. A QA team may decide to concentrate testing on the devices and journeys historically associated with the greatest risk.
None of this is unusual. Product teams cannot wait for perfect information before acting.
The problem begins when assumptions gradually become indistinguishable from facts.
Once an idea becomes an epic, receives a priority and is assigned to a quarter, it acquires a certain permanence. The team sees the work that needs to be delivered, but the evidence that made that work important becomes less visible.
That is particularly problematic because a backlog was never meant to be static. Atlassian describes the product backlog as a “live” artefact that should be updated as new information becomes available. Its guidance on backlog refinement makes the same point: refinement is an ongoing process in which priorities are reviewed against feedback and business value.
In other words, changing the backlog when the evidence changes is not a deviation from good product management. It is part of it.
The same principle applies beyond Agile methodologies. In Harvard Business Review, Harvard Business School professor Stefan Thomke and Gary Loveman argue for a more scientific approach to management: challenge assumptions, formulate hypotheses, experiment and allow evidence to change the decision.
A completed ticket tells us that something was delivered. A released feature tells us that something reached users.
Neither, by itself, tells us whether the original problem was correctly understood or whether the expected outcome was achieved.
A few weeks away from the usual operating rhythm may not seem enough to invalidate a strategy.
Often, they aren't.
But the roadmap is not the only thing that matters.
While priorities remain on a planning board, users continue interacting with the product. New support tickets accumulate. Analytics collect additional behavioural data. Releases introduce changes. Competitors launch alternatives. Acquisition sources shift. New operating-system versions appear. Business priorities move.
The important point is not that everything has changed.
It is that teams should know what has changed before assuming nothing has.
This is particularly relevant for digital experiences because the product itself is only one part of what determines the experience.
A checkout can be technically identical in June and September while the people using it, their expectations, traffic sources or device mix have changed. A journey that performed well under one set of conditions may behave differently under another.
As we explored when looking at how products perform in real-world conditions, users rarely interact with digital services in the controlled environment in which they were designed or tested. They switch networks and devices, multitask, get interrupted and encounter conditions that are difficult to reproduce internally.
Summer makes some of these changes particularly visible, but the principle applies throughout the year: context is part of the experience.
For teams returning to Q4 planning, the implication is straightforward. Before checking the status of the work, it is worth checking the status of the evidence behind it.
Suppose a team has prioritised a checkout redesign because research conducted nine months ago identified friction during payment.
That evidence is considerably stronger than an internal opinion.
But is it still enough?
Has the checkout changed since the research was conducted? Has the payment mix changed? Are the same problems appearing in analytics or customer support today? Does the issue still occur across the devices and environments that matter most?
Evidence does not suddenly become useless because it is six or twelve months old. Its relevance depends on how much the product, audience and surrounding context have changed since it was collected.
A study of a relatively stable enterprise workflow may remain useful for a long time. Findings about a fast-moving consumer journey may require much more frequent verification.
The question, therefore, is not simply “Do we have data?” but “Is this still the right evidence for the decision we are making today?”
That evidence might be recent and directly observed through behavioural data, user research or testing. It might be historical but still relevant. It might be indirect, coming from support conversations, stakeholder feedback or market signals. Or a closer look may reveal that something everyone assumed to be known has never actually been validated.
Making those distinctions visible can change a prioritisation conversation considerably.
Most roadmaps already contain some form of prioritisation based on impact, effort, urgency or strategic value.
What they don't always make visible is uncertainty.
Imagine a team entering September with three major Q4 priorities: improving onboarding, redesigning part of the checkout and developing a feature repeatedly requested by customers.
All three may sound reasonable. The usual conversation will focus on resources, expected impact and delivery dates.
But before ranking them, there is another question worth asking:
What needs to be true for each initiative to work?
Take onboarding. The working assumption might be that users abandon the process because there are too many steps.
It sounds plausible. It may even be correct.
But the number of steps might not be the problem at all. Users may not understand why certain information is required. The friction could appear earlier in the journey. Or people may simply not perceive enough value to complete the process.
A team could successfully shorten onboarding and still fail to improve completion.
This is where uncertainty becomes useful as a prioritisation criterion.
Instead of looking only at expected impact and effort, teams can consider two additional variables: how confident are we in the assumption, and what would happen if we were wrong?
An assumption with limited evidence but little consequence may be perfectly acceptable.
An important assumption supported by strong, recent evidence may not need another round of research.
The interesting area is where low confidence meets high consequence.
For Product, that might be the belief that a requested feature addresses the actual customer need.
For UX, it could be the assumption that a behaviour observed in previous research still represents today's users.
For QA, it might concern which workflows, integrations, environments or devices carry the greatest risk ahead of a major release.
In these cases, research and testing stop being activities performed only after a solution has been designed or built. They become tools for reducing uncertainty before the expensive decisions are made.
There is another reason to revisit evidence before Q4 execution accelerates: it is often fragmented across the organisation.
Product may be looking at adoption and business metrics. UX may have qualitative evidence from research. QA may see recurring defects and edge cases that aren't visible in product analytics. Customer Support knows which problems repeatedly generate frustration. Sales hears objections and requests that may never appear in a dashboard.
Each source describes a different part of reality.
Consider a product journey with a declining conversion rate.
Analytics can show where users leave, but not necessarily explain why. User research can reveal confusion, expectations or motivations, but a qualitative study alone cannot establish how widespread a behaviour is. QA testing may uncover a functional issue affecting a specific environment without revealing its overall commercial impact.
Looking at these signals together creates a more complete picture than expecting any single source to provide the answer.
There is also an organisational benefit: evidence gives cross-functional teams something concrete around which to align.
The effect can be seen in the Lottomatica case, where UX, Product and Market Research worked iteratively with users, moving from an initial usability test to design changes and subsequent validation. The research helped the team confirm some assumptions while identifying issues in areas that had not initially seemed critical.
This is an important distinction.
Evidence is valuable not only because it tells teams something about users. It also gives different functions a common reference point for making decisions.
“Let's validate it” can sound like adding another phase to an already crowded roadmap.
It doesn't have to.
The method should match the uncertainty.
If the question concerns behaviour, existing analytics may already contain part of the answer. If the team needs to understand motivations or expectations, qualitative research may be more appropriate. If the uncertainty concerns comprehension or interaction, a prototype can be tested before development begins. If the risk is technical, exploratory, compatibility or real-world testing can expose problems that controlled internal environments fail to reproduce.
Sometimes apparently small interactions justify investigation because their business or user impact is disproportionately large.
In the Nexi Mobile POS project, for example, user research investigated specific aspects of interaction with a numerical keypad, including the visibility and comprehension of individual controls. The size of an interface element does not necessarily reflect the impact it can have on the overall experience.
And sometimes testing simply confirms that the team's original assumption was correct.
That is useful too.
When an initiative is about to absorb significant development time or influence an important business outcome, greater confidence in the underlying decision has value in itself.
There is one final risk worth considering before Q4 begins: evidence does not remain equally useful over time.
A product team can make a perfectly reasonable decision based on research, analytics or testing and still need to revisit it months later. The issue is not whether the original evidence was good, but whether the conditions in which it was collected are still comparable to today's.
A significant product release can change user behaviour. A new acquisition channel can bring in a different audience. Changes to pricing, authentication or payment methods can alter an established journey. New devices, operating-system updates and third-party integrations can introduce technical conditions that did not exist when a testing strategy was defined.
This is why evidence-led decision-making should not be confused with simply having data. What matters is whether the available evidence is recent, relevant and appropriate to the decision being made.
For teams entering Q4, that creates a useful distinction between decisions that need to be made again and decisions that simply need to be checked.
Most priorities will probably survive that check. Some may need stronger evidence. A few may turn out to be solving a problem that has changed, disappeared or was never fully understood in the first place.
Finding those few is where the value lies.
None of this means spending the first weeks of September reopening every decision made before summer. That would create another kind of inefficiency.
A more useful approach is to focus on the initiatives with the greatest combination of business impact, uncertainty and cost of being wrong.
For each major Q4 priority, teams should be able to explain four things:
What problem are we trying to solve?
What evidence tells us that the problem still exists?
Which assumptions does our proposed solution depend on?
What new information would materially change our decision?
If those answers are clear, the backlog can do exactly what it is supposed to do: turn priorities into action.
If they are not, moving a ticket into In Progress will not make the underlying uncertainty disappear.
And this is perhaps the most useful way to think about the September restart. The objective is not to slow delivery down or question the roadmap for the sake of it. It is to make sure that the final and often most intense part of the year begins with priorities that still reflect the reality outside the organisation.
Because a backlog is useful for remembering what a team planned to do.
Evidence is what tells you whether it is still the right thing to do.
Have assumptions you need to validate before accelerating your Q4 priorities? With UNGUESS, you can put them in front of real users and test your products in real-world conditions before investing time and resources in the wrong direction.