Solutions / Custom Services / Product Testing
Product Testing
Internal testing tells you the product works as designed. User testing tells you whether it works as used, which is a different question.
Real users, real conditionsTested as used rather than as designed, which is a different and more useful question.
Against the real alternativesComparison with the products a buyer would genuinely choose between.
Observed, not self-reportedTask observation surfaces failures that a questionnaire never captures.
In simple terms
You tell us what you are building or launching. We put it in front of real users under realistic conditions and tell you where it works and where it does not.
What Product Testing is. Product testing programmes cover concept, prototype, usability, sensory and in-home testing with recruited users matched to your target, measuring performance against the alternatives they would otherwise choose.
What makes this different. Fieldwork quality is decided long before anyone answers a question: in the screener, the instrument design and the quality control. That is where most of our effort goes, and it is why the findings hold up.
What the service covers
These are the areas we work across, such as the ones below. We combine them in the proportion your question needs, and we tell you which ones your question does not need.
Concept testing
Reactions to positioning and concept before build, with comparison against the alternatives in the user's real choice set.
Usability testing
Observed task completion with think-aloud protocol, which surfaces failures self-report never captures.
Comparative testing
Blind or branded head-to-head against competitor products under controlled conditions.
In-home and in-use
Extended testing in real conditions, where problems appear that never show in a controlled session.
Questions this answers
If something like this is on your agenda, the engagement is already half scoped.
- Does this work for real users, or only for us?
- How does it perform against the competitor product side by side?
- Which feature do users actually value, and which do they ignore?
- Where does the product fail in ordinary use?
- Is the concept worth building at all?
How you can use this
A few of the situations where this service does real work.
Validating before build
A concept is approved and untested with anyone outside the company.
Users failing at something
Support volume points at a problem you cannot reproduce.
Comparing with a competitor
You believe you are better and cannot demonstrate it.
Choosing between features
Several options compete for the same development slot.
What changes for your business
The practical difference between running on this and running on what you have now, such as the following.
You get evidence where none was published
Primary research reaches the questions no dataset answers, which is most of the questions that actually matter.
The data is clean enough to rely on
Screening, quality control and documented exclusion rules mean findings survive scrutiny rather than collapsing under it.
You get topline fast
Headline findings arrive within days of fieldwork closing, so decisions do not wait on the full report.
Respondents tell you the truth
Independent fieldwork changes what people are willing to say, particularly about what is going wrong.
You can run it again
Instrument, screener and methodology are handed over, so tracking the same question over time costs a fraction of the first wave.
You test before you commit
Concepts, prices and propositions get checked with real respondents while changing them is still cheap.
Who this is for
Roles that commission this work most often include those below. Each asks a different question and gets a different cut of the same evidence.
Research and insight teams
We need fieldwork we cannot run in-house.
Specialist capacity to your standards, with instrument, dataset and methodology handed over for reuse.
Product and marketing leadership
What do our customers actually think, rather than what we assume?
Evidence from real respondents, with verbatims, so findings can be acted on rather than debated.
Strategy and commercial leadership
Is there enough here to justify the investment?
Primary evidence on demand, pricing and competitive position where no published data exists.
Operations and quality leadership
Is what we designed actually happening in practice?
Independent observation and measurement, reported by location and behaviour so it supports coaching.
Investors and diligence teams
Can we verify this independently before we commit?
Primary market and customer evidence gathered to a defined protocol, with full quality control.
Typical clients
How we work
The third step is the one that makes the output usable, and it is the one most work of this kind skips.
Understand the decision
We start from what the findings have to support, because that decides sample, method and the precision you actually need to buy.
Design and field properly
Screener, instrument and protocol built and piloted before full fieldwork, with quality control running throughout rather than checked at the end.
Read it for your position
Findings are interpreted against your market position and the decision in front of you, not reported as a neutral data dump.
Deliver so it can be reused
Dataset, verbatims, instrument and methodology handed over, so the next wave costs a fraction of the first.
What you receive
Full dataset with respondent metadata and quality flags
Topline findings within days of fieldwork close
Analytical report with verbatimscross-tabs and implications
Instrumentscreener and methodology note for reuse
Access to Phi, our AI research platformincluded
Every engagement comes with access to Phi. Ask questions of your own findings in plain language, pull the evidence behind any number, and keep querying long after the work is delivered. Your team gets the working intelligence, not just the document.
Open PhiWays to start
Tell us the decision and the date it has to be made by. We will recommend the smallest engagement that gets you there.
Pilot wave
A small first wave to validate approach, screener and instrument.
Full fieldwork
Complete fieldwork to target quota, with analysis and reporting.
Tracking programme
The same instrument repeated on a set cycle, reported as change over time.
Common questions
How many users do we need?
Usability issues surface with five to eight users per segment. Comparative and sensory testing need larger samples for statistical comparison.
Blind or branded?
Both, where budget allows. The gap between blind and branded results tells you what your brand is actually worth in the category.
Who owns the work?
You do. It is exclusive to you, it is not resold, and working files are handed over unlocked.
Will you sign an NDA?
Yes, before the first conversation if you prefer.
Tell us the decision you are facing
Send us the question, the context and your timeline. You get back a recommended first step, what evidence it needs, and what it costs. Not a capability deck.
