Solutions / Custom Services / Product Testing

Product Testing

Internal testing tells you the product works as designed. User testing tells you whether it works as used, which is a different question.

Real users, real conditionsTested as used rather than as designed, which is a different and more useful question.

Against the real alternativesComparison with the products a buyer would genuinely choose between.

Observed, not self-reportedTask observation surfaces failures that a questionnaire never captures.

depth reach InterviewsPanelsSurveysField a mix, weighted to the question
Qualitative depthQuantitative reach

In simple terms

You tell us what you are building or launching. We put it in front of real users under realistic conditions and tell you where it works and where it does not.

What Product Testing is. Product testing programmes cover concept, prototype, usability, sensory and in-home testing with recruited users matched to your target, measuring performance against the alternatives they would otherwise choose.

What makes this different. Fieldwork quality is decided long before anyone answers a question: in the screener, the instrument design and the quality control. That is where most of our effort goes, and it is why the findings hold up.

What the service covers

These are the areas we work across, such as the ones below. We combine them in the proportion your question needs, and we tell you which ones your question does not need.

Concept testing

Reactions to positioning and concept before build, with comparison against the alternatives in the user's real choice set.

Usability testing

Observed task completion with think-aloud protocol, which surfaces failures self-report never captures.

Comparative testing

Blind or branded head-to-head against competitor products under controlled conditions.

In-home and in-use

Extended testing in real conditions, where problems appear that never show in a controlled session.

Questions this answers

If something like this is on your agenda, the engagement is already half scoped.

  • Does this work for real users, or only for us?
  • How does it perform against the competitor product side by side?
  • Which feature do users actually value, and which do they ignore?
  • Where does the product fail in ordinary use?
  • Is the concept worth building at all?

How you can use this

A few of the situations where this service does real work.

Validating before build

A concept is approved and untested with anyone outside the company.

An honest read before the development spend

Users failing at something

Support volume points at a problem you cannot reproduce.

The failure observed and located precisely

Comparing with a competitor

You believe you are better and cannot demonstrate it.

Head-to-head evidence under controlled conditions

Choosing between features

Several options compete for the same development slot.

Measured preference rather than internal advocacy

What changes for your business

The practical difference between running on this and running on what you have now, such as the following.

You get evidence where none was published

Primary research reaches the questions no dataset answers, which is most of the questions that actually matter.

The data is clean enough to rely on

Screening, quality control and documented exclusion rules mean findings survive scrutiny rather than collapsing under it.

You get topline fast

Headline findings arrive within days of fieldwork closing, so decisions do not wait on the full report.

Respondents tell you the truth

Independent fieldwork changes what people are willing to say, particularly about what is going wrong.

You can run it again

Instrument, screener and methodology are handed over, so tracking the same question over time costs a fraction of the first wave.

You test before you commit

Concepts, prices and propositions get checked with real respondents while changing them is still cheap.

Who this is for

Roles that commission this work most often include those below. Each asks a different question and gets a different cut of the same evidence.

Research and insight teams

We need fieldwork we cannot run in-house.

Specialist capacity to your standards, with instrument, dataset and methodology handed over for reuse.

Product and marketing leadership

What do our customers actually think, rather than what we assume?

Evidence from real respondents, with verbatims, so findings can be acted on rather than debated.

Strategy and commercial leadership

Is there enough here to justify the investment?

Primary evidence on demand, pricing and competitive position where no published data exists.

Operations and quality leadership

Is what we designed actually happening in practice?

Independent observation and measurement, reported by location and behaviour so it supports coaching.

Investors and diligence teams

Can we verify this independently before we commit?

Primary market and customer evidence gathered to a defined protocol, with full quality control.

Typical clients

Enterprises needing evidence for a specific decisionMid-size companies without fieldwork capabilityIn-house insight teams needing extra capacityInvestors validating a market or customer baseOrganisations researching unfamiliar markets

How we work

The third step is the one that makes the output usable, and it is the one most work of this kind skips.

Step 1

Understand the decision

We start from what the findings have to support, because that decides sample, method and the precision you actually need to buy.

Step 2

Design and field properly

Screener, instrument and protocol built and piloted before full fieldwork, with quality control running throughout rather than checked at the end.

Step 3

Read it for your position

Findings are interpreted against your market position and the decision in front of you, not reported as a neutral data dump.

Step 4

Deliver so it can be reused

Dataset, verbatims, instrument and methodology handed over, so the next wave costs a fraction of the first.

One finding established once Upstream supplierProducer or provider Channel or partnerInvestor Secure capability earlyBring the change forward Prepare for the shiftReprice the exposure four different recommended actions
Step three in practice. Your position decides what a finding means.

What you receive

Full dataset with respondent metadata and quality flags

Topline findings within days of fieldwork close

Analytical report with verbatimscross-tabs and implications

Instrumentscreener and methodology note for reuse

Access to Phi, our AI research platformincluded

Every engagement comes with access to Phi. Ask questions of your own findings in plain language, pull the evidence behind any number, and keep querying long after the work is delivered. Your team gets the working intelligence, not just the document.

Open Phi
Request a sample Real work with the client's identity removed.

Ways to start

Tell us the decision and the date it has to be made by. We will recommend the smallest engagement that gets you there.

1-2 weeks

Pilot wave

A small first wave to validate approach, screener and instrument.

3-6 weeks

Full fieldwork

Complete fieldwork to target quota, with analysis and reporting.

Quarterly

Tracking programme

The same instrument repeated on a set cycle, reported as change over time.

Common questions

How many users do we need?

Usability issues surface with five to eight users per segment. Comparative and sensory testing need larger samples for statistical comparison.

Blind or branded?

Both, where budget allows. The gap between blind and branded results tells you what your brand is actually worth in the category.

Who owns the work?

You do. It is exclusive to you, it is not resold, and working files are handed over unlocked.

Will you sign an NDA?

Yes, before the first conversation if you prefer.

Tell us the decision you are facing

Send us the question, the context and your timeline. You get back a recommended first step, what evidence it needs, and what it costs. Not a capability deck.