top of page

Sample Size in Sensory: Is BIGGER Always Better?

  • 11 hours ago
  • 2 min read


Imagine this situation:


R&D reformulates its company’s flagship barbecue sauce - the business’ cash cow. The reformulation delivers sizable 𝗰𝗼𝘀𝘁 𝘀𝗮𝘃𝗶𝗻𝗴𝘀, but management refuses to 𝙧𝙞𝙨𝙠 𝙘𝙤𝙣𝙨𝙪𝙢𝙚𝙧𝙨 rejecting the product.


The sensory program established a robust 𝗿𝗶𝘀𝗸 𝗽𝗿𝗼𝗳𝗶𝗹𝗲 based on suitable internal vs. consumer research for their specific product category:


    • Tetrad test

    • With N = 47

    • Giving 𝟵𝟬% 𝗽𝗼𝘄𝗲𝗿

    • To detect a consumer threshold of 𝘋𝘦𝘭𝘵𝘢𝘙 = 1.20

    • At alpha = 0.05


Driven by brand anxiety, management insists on DOUBLING the sample size to "minimize the risk of missing a difference."


The outcome:


   • Method: Tetrad

   • Sample size: 94 panelists

   • Results: 42 correct responses

   • 𝘱 = 0.014 (𝘱 < 0.05) ➔ Reformulation REJECTED! ❌


Was this the right decision? Or did R&D just throw away a great cost-saving opportunity? 💸


Here is what really happened behind the scenes:


    • The estimated 𝘥’ value was 0.82 - lower than the 𝗰𝗼𝗻𝘀𝘂𝗺𝗲𝗿 𝘁𝗵𝗿𝗲𝘀𝗵𝗼𝗹𝗱 of 1.20. 

    • Because the sample size doubled, the statistical power shot up. ↗️

    • The standard difference test became so sensitive that it flagged a real, but 𝘱𝘳𝘢𝘤𝘵𝘪𝘤𝘢𝘭𝘭𝘺 𝘪𝘯𝘴𝘪𝘨𝘯𝘪𝘧𝘪𝘤𝘢𝘯𝘵, sensory difference.


💡 𝙒𝙝𝙚𝙧𝙚 𝙙𝙞𝙙 𝙩𝙝𝙚 𝙧𝙚𝙨𝙚𝙖𝙧𝙘𝙝 𝙜𝙤 𝙬𝙧𝙤𝙣𝙜?


They kept evaluating the data with a standard 5% difference test instead of updating the framework to a similarity (equivalence) test.


When running a direct similarity test against the consumer threshold (𝘋𝘦𝘭𝘵𝘢𝘙 = 1.20), the conclusion flips: 𝘥' = 0.82 is statistically LOWER than 1.20 at the 95% confidence level. The products are essentially equivalent to the consumer base! 📏


To avoid abandoning viable innovations when sample sizes increase, sensory programs must shift from simple difference testing to direct similarity testing against known consumer thresholds. Yet, this is still surprisingly rare in CPG sensory programs.


So, is bigger always better for sample size in sensory? 𝗡𝗢, 𝗶𝗻𝗱𝗲𝗲𝗱! Without aligning a statistical method to the sample size, 𝗕𝗜𝗚𝗚𝗘𝗥 will just result in detecting irrelevant differences and rendering successful reformulations much more difficult to achieve. 💸


All of these insights were only possible because the company originally invested in understanding their true consumer threshold (𝘋𝘦𝘭𝘵𝘢𝘙 = 1.20). Without that benchmark, they would be flying blind. 🫣


What is your experience with internal sensory programs?


Are they usually equipped to run direct similarity testing when sample sizes shift, or do they still rely solely on 𝘱 < 0.05 difference tests?



 
 

DAVIS SENSORY INSTITUTE .LLC

Contact

dsiteam@davissensory.com
(530) 750-6503

946 Olive Drive, Suite 9
Davis, CA 95616

Hours

Monday - Friday
10:00 am - 6:00 pm

Note: Consumer studies may affect business hours.

  • Instagram
  • LinkedIn

© Davis Sensory Institute 2026

Architecture Design Credits: Maria

Selected Photo Credits: Julia Ogrydziak

bottom of page