ProductForce

Research Synthesis Hero · Product Management · Estimated 7 min read

Team

Mona Saeed (Data Scientist),
Harry Liu (SWE)

Role

AI Product Manager · UX

Stack

AI/ML: Gemini 2.5 · RAS · Hierarchical clustering

App Deployment: Next.js · FastAPI · Supabase Vector · Cloud Run

Project Highlight

The Superpower for Product Teams

ProductForce fuels innovation economy and collaboration by giving every team member the superpower to turn a pile of raw interviews into a prioritized, evidence-backed list of actionable items. What used to take hours for complex pattern recognition across a large quantity of interviews now takes minutes, with human control at every step of the decision-making process.

Background

Client Background

Navi is an early-stage startup that builds products that help people and organizations earn skills to unearth, articulate, and solve problems. Their current offering includes 2 tools: Problem Formula, an AI-assisted flow that walks users step by step through documenting and refining a problem statement, and Behavioral Map, an iterative pass at identifying key behaviors and pain points.

User interviews are the fastest and cheapest way to start and validate a process. ProductForce sits upstream of everything Navi already sells, providing a centralized repository for the team to collect interview insights that feed directly into Problem Formula and the value transaction map.

Navi's Problem Formula tool with its AI assistant Interview documentation tracked in Linear Hand-drawn scene of the bootcamp: the big screen, the two of us presenting, and the behaviour map
At the bootcamp: Problem Formula on the big screen, the behaviour map re-drawn by hand, and every interview typed up in Linear
Problem

Insight Debt

Product managers and UXRs spend 12–15 hours synthesizing every five interviews. That manual burden creates insight debt: patterns lost to fatigue, or buried in static documents nobody reopens. Roadmaps end up decoupled from the raw user evidence that should support them.

80%

of research time lost to clerical tagging

40%

of critical insights missed due to fatigue

Customer Discovery

Product-Market Fit

We identified Life and Health Sciences as the initial target market, which is the next industry Navi plans to move into. However, we discovered that many internal decisions for R&D innovation bypass interviews and are dependent on organizational structures.

We also started this pursuit to create a "teleprompter app" to assist with user interviews. After running over twenty live interviews across the product and healthcare ecosystem, we realized we needed to adapt our target market for this project. We narrowed down who the product serves and which key behavior of the interview process to own.

A live interview: the video call with the team's faces alongside the write-up in Linear 20+ LIVE INTERVIEWS · THE EVIDENCE BASE THE INTERVIEWS CONVERGED HERE Product Managers 6 INTERVIEWED “I want the synthesis done before the roadmap meeting, not after it.” Live notes fail mid-conversation Confirmation bias Synthesis lands after the roadmap decision Healthcare Professionals 2 INTERVIEWED No structured path to advocate change IRB friction NOT THE WEDGE NOT THE WEDGE Duke Health Students 2 INTERVIEWED Objectives are the hard part
It is a slow, hierarchical process to push for R&D innovation and change in mindset in Health Sciences, because many internal decisions depend on organizational structures even with carefully prioritized insights.The interviews converged on one persona and one phase: the PM/UXR synthesizing post-interview at volume.
The persona we narrowed to: Jordan Adams, Product Manager / UXR in a tech company
Feature Prioritization

From Persona to Backlog

At the end of Sprint 1, my team and I brainstormed and sorted features by their impact on Jordan's job and our team's technical feasibility in Sprint 2. We validated these features with previous interviewees who match the target persona and decided to focus on team collaboration, HITL control of synthesis results, and insight prioritization.

The Sprint 2 backlog, try dragging the cards!
FEATURE PRIORITIZATION

Technical Trade-offs

Summaries were untrustworthy without clear labels: nobody could tell which sentence came from a user and which came from the model. We rebuilt the analysis layer as retrieval-augmented synthesis (RAS) so that every claim is forced to cite the transcript span it came from, and any claim that cannot cite is dropped before it renders. Every piece of text inside ProductForce is editable for human control.

Pattern discovery needs scale to tune, and specificity to trust. We calibrated clustering heuristics on high-volume Amazon review data, then tested on specialised ServiceNow solutions-engineer transcripts to:

Gemini 2.5 Pro · 1M context

Reasons across 50+ interviews at once, finding patterns that span sessions instead of summarizing them one at a time.

Hierarchical clustering

Unsupervised NLU groups repeated pain points into tiered themes, cutting manual tagging burden by ~90%.

RAS grounding

Zero-hallucination policy: no summary is viewable without its source quote attached and highlighted.

WORKING HERO

ProductForce Demo

The 5-step interaction flow (prototyped with Vercel V0)

Ingest
Upload transcript
Supports .vtt and .txt · scrubbed on device before upload
Drop a transcript file…
Names redacted Emails redacted Account IDs redacted
Analyzing · servicenow-se-04.vtt
Gemini 2.5 · five-step NLU pass
  • Extract candidate pain points14s
  • Cluster into tiered themes39s
  • Link each theme to source excerpts71s
  • Detect emotion peaks104s
  • Score confidence & drop ungrounded claims128s
Verify the insight against its source
No summary is viewable without the quote that produced it
AI insight

Solutions engineers rebuild the same demo environment for every customer call, losing prep time to setup rather than tailoring.

Source verified
Transcript · 00:19:42

“Honestly the demo itself is fine. It's that I rebuild the whole environment from scratch every single time — that's the part eating my week.”

Weighted priority
Auto-suggested from four inputs — the PM can override any of them
Customer impact8.6
Strategy fit7.4
ARR exposure9.1
Trend velocity6.2
Weighted score0.0P1 — this quarter
Export to the backlog
The insight leaves as a ticket, not a slide
JiraLinear
PF-214 · created
Persist reusable demo environments per account
P1Score 8.14 interviews citedGrounded ✓

a. The dashboard below shows project cards, upload entry point, and interview table. Interviews are sorted by projects and time conducted.

b. The project prioritization page: every verified problem area across the project's interviews, filterable by persona, priority, and date. Clicking a row opens a detail panel with supporting excerpts and source transcripts.

Evaluation

Success Metrics

Every claim about the model was tied to a KPI with a baseline. Clustering quality improved from a Sprint 1 silhouette score of 0.72.

MetricDefinitionResult
Silhouette scoreMathematical distinctness of thematic clusters0.78
Recall rateNuanced patterns identified vs. human coding+50% signal
Grounding rateInsights directly verifiable by transcript quotes100%
API latencyEnd-to-end processing of a 60-minute session~128 sec

Impact on analysis velocity. In a pilot with ServiceNow solutions engineers, synthesis time per five interviews collapsed by 88%, reclaiming strategic time.

Manual synthesis
15.0 hours
With ProductForce
1.8 hours
Responsible AI

Trust Decisions I Owned as PM

PII redaction

A local scrubbing layer strips customer names and emails before data reaches the LLM API. In enterprise B2B, security is a sales enabler, not a checkbox.

Hallucination defense

The UI forces quote-back verification: a summary cannot be viewed without its original transcript evidence attached.

Bias mitigation

Hierarchical clustering protects quiet voices so niche feedback isn't swallowed by majority sentiment.

Residual risk

Output quality still depends on transcript quality; noisy audio files carry a low-fidelity warning rather than a silent degradation.

BUSINESS ADOPTION POTENTIAL

~$0.15 Per Synthesis Cycle

Built on Next.js, FastAPI, and a Supabase vector database, deployed to Google Cloud Run so research spikes scale serverlessly without idle overhead. At roughly fifteen cents per synthesis cycle, the unit economics support a seat-based SaaS model aimed at mid-market and enterprise product ops. Estimated adoption: 850k product teams x $1.2k/yr = $1.02B TAM.

$0.15

processing cost per synthesis cycle

SaaS product teams (45%) Enterprise UX orgs (35%) Growth startups (20%)
850k

product teams in the beachhead segment

FUTURE SCOPE

Workspace Integration

Q3 2026

Native Zendesk and Salesforce sync

Q4 2026

AI video highlight reels

Olá!
Q1 2027

Cross-language synthesis

Q2 2027

Predictive churn detection

LESSONS LEARNED

Design an AI System

Researchers want AI to prepare the work

Not to do it. The moment the model made the final call, trust collapsed.

Ethics is a feature

PII scrubbing was the single most requested capability in enterprise conversations, ahead of any analysis improvement.

Calibration needs volume

Pattern-discovery sensitivity can't be tuned on your own small corpus; borrowing a massive public dataset made the specialised one usable.

Never too late to pivot

Our pivot from developing a tool for Health Sciences to broadening our scope to universal teams changed Navi founders' perspectives on leveraging interviews as their strong suit for innovation.

Harry and me presenting ProductForce at the Duke CFCI Product Demo Day
P.S. me and my teammate Harry at Duke CFCI Product Demo Day

Thanks for reading :)

More projects

Onboarding Agent

Internship with Pinecone · UX Research · PRD · MVP

Gradually

Product Design · Wellness Plug-In