
AskGauge
A conversational AI for TrustGauge, a nonprofit that has tested consumer products for nine decades. AskGauge turns that archive of ratings, reviews, and articles into plain answers and personalized recommendations, in the voice of a TrustGauge expert, in a few seconds.
Business Overview
TrustGauge is a US nonprofit that has tested and rated consumer products for close to ninety years. Shoppers trust it because it buys what it tests, takes no advertising, and publishes the results without favor. That trust is its whole business, and it sits on an archive of ratings, lab data, and articles that no retailer can match.
The organization's innovation team wanted to put that archive one question away. Imagine typing "I am tall and sleep on my back, which mattress suits me?" and getting a real recommendation grounded in TrustGauge's own testing. They brought in Xpiderz as their development partner for a seven-month build, working as one team with TrustGauge's engineers and product managers to design, build, and launch AskGauge to members in a closed beta.
Challenge
A shopper should be able to ask AskGauge a question the way they would ask a knowledgeable friend and get a personal recommendation in the voice of a TrustGauge expert. That sentence hides most of the difficulty. Real questions are vague, mix several needs at once, and use words the lab data never uses.
The system had to work out what a person actually meant, whether they were doing general research, asking about one model, or stuck between two choices, then find the right ratings, articles, and specs across hundreds of product categories and turn them into an answer that is correct, useful, and sounds like TrustGauge. It also had to stay safe in front of the public, refuse harmful requests, and never invent a rating that does not exist.
Solution
We built AskGauge as a retrieval augmented generation system driven by a custom agent rather than a single prompt. Every question passes through a set of configurable steps: guardrails screen it, a refiner works out the topic and intent, a retrieval layer pulls the matching products, articles, and ratings, and a synthesis step writes the answer in TrustGauge's tone with links to the source pages.
The first version used a general-purpose agent that was easy to maintain but not accurate enough. We replaced it with purpose-built subsystems that can be tuned independently, then ran them in parallel so a fully researched answer still arrives in seconds rather than minutes. Shoppers can ask about a category or a specific model on web and mobile, and every answer links back to the TrustGauge article or rating behind it.
AskGauge shipped with six working parts:
-
Ground Truth Defined With Experts
Before writing a prompt, we explored the structure of the archive and worked with TrustGauge's product testers to define the expected answer to a bank of real questions. That set became the yardstick for every later change, so we always knew whether the system was getting better or worse.
-
Guardrails and an Evaluation Suite
Layered moderation screens for hate speech, self-harm, sexual content, and violence before a question reaches the model. An automated evaluation suite, plus red-team testing, reruns the full question bank on every change. Guardrail performance improved more than tenfold over the project.
-
Query Refinement and Routing
A refiner rewrites the shopper's question and identifies topic and intent, then routes it to the right part of the data. It knows the difference between general research, a question about one model, and a shopper weighing two options, and it learned that from a handful of examples.
-
Teaching the Model to Read Lab Data
Ratings for hundreds of categories all measure different things. We built a data pipeline that generates a schema per category and, with TrustGauge's experts, wrote instructions explaining what each feature means to a buyer, so the model interprets a humidistat accuracy score the way a human tester would.
-
Memory Across the Conversation
Language models forget the last question on their own. AskGauge started with a short buffer of recent turns and ended with a hybrid memory that keeps both the latest exchanges and long-term context, so "are any of those latex-free?" works after a mattress recommendation.
-
Fast Answers Under Real Load
Retrieval and summarization subsystems run in parallel, keeping responses to a few seconds even as accuracy rose. TrustGauge staff and product testers stress-tested the system ahead of launch to confirm it could carry the concurrent load of the beta audience.
Measured Results
The numbers speak
What the seven-month build delivered by the time AskGauge reached beta members.
One Team, Two Logos
This was a joint engineering effort, not a handoff. Xpiderz engineers and TrustGauge engineers reviewed each other's code, ran shared stand-ups, and met several times a week with the innovation lab's leadership, product manager, and project manager to plan research, testing, and delivery. GitHub, Jira, and Slack were the shared workspace.
Design ran the same way. We held review sessions with TrustGauge's design team so AskGauge matched the rest of the brand, and worked with its user research team on rounds of external feedback that shaped both the interface and the model's behavior.
Outcome
AskGauge launched to TrustGauge members as an invitation-only beta on web and mobile browsers:
- Real recommendations, not search results. A member describes their situation in plain words and gets a specific pick, the reasoning behind it, and links to the ratings and articles that back it up.
- Trust preserved. Every answer is grounded in TrustGauge's own testing. If the archive does not cover a product, AskGauge says so rather than guessing, and harmful or off-topic requests are declined.
- A foundation that can grow. Coverage is limited to categories TrustGauge experts have examined, and the schema pipeline means new categories can be added without rebuilding the system.
Next Steps
The rollout is deliberately gradual. New members come off the waitlist in waves, each wave feeds the evaluation suite with fresh questions, and the joint team reviews the weak spots before opening the next. Category coverage expands as TrustGauge's testers examine new products, and the same pipeline that reads a mattress rating is already reading cars, appliances, and electronics.
We have spent ninety years earning the right to tell people what to buy. Handing that voice to an AI was the scary part. Xpiderz treated our testers as the source of truth, measured every change against it, and built something our members can ask a real question and trust the answer.
Head of Innovation Lab
TrustGauge
















