Product Strategy — Demand, ICP & Competitive Positioning
Purpose. Answers the three questions the rest of our docs never answer: What problem do we solve and is the demand real? Who is our customer and why do they pay? What is the competitive landscape and why do we win? It complements
PRODUCT_OVERVIEW.md(what/how/when) with the why-we-win layer.Honesty markers. This analysis is built on an evidence base of one real paying engagement (Kaito) plus internal narrative. Claims are labeled by evidence strength throughout. Confidence should rise or fall as customers 2, 3, 4 land. Competitive pricing data was researched 2026-07 and goes stale fast.
中文版:
PRODUCT_STRATEGY.zh.mdLast reviewed: 2026-07-04.
1. The Problem We Solve — and Whether the Demand Is Real
The job, in one sentence
"I have high-volume, templated, externally-facing communication work that drains senior time — but a single mistake is expensive, so I can't hand it to plain outsourcing or unsupervised AI."
We call this "guard-railed labor arbitrage." It is one job with two faces, and both must hold for h.work to be the right hire:
- the arbitrage face: the work is repetitive and doesn't justify a senior salary (or any dedicated headcount);
- the guard-rail face: the work carries brand, compliance, or regulatory risk, so "cheap" alone is disqualifying.
Products that serve only one face — pure-AI agents (arbitrage without accountability) or BPO (labor without AI leverage or engineered guardrails) — structurally miss this job.
Demand validation — two hypotheses, tested separately
H1 — Labor arbitrage demand ("companies will pay a monthly fee to offload templated ops work to a managed AI+Expert employee")
| Sub-hypothesis | Confidence | Evidence | What would falsify it |
|---|---|---|---|
| The work exists at volume | High | Kaito engagement docs: 5–7 posts/campaign × 2–4 campaigns/mo + 48–72h reply triage; "drains senior time without justifying senior salary" | — |
| Customers pay for it (vs. ChatGPT + an intern) | Medium | Kaito signed — but n=1 | Kaito churns; prospects 2–3 don't close |
| Price > our delivery cost (Expert time) | Low — most dangerous | None. At autoRespondThreshold=101 every reply burns Expert time; ~80% of Kaito triage replies are FAQ-shaped | Unit economics show negative or BPO-level gross margin |
| The job needs persona + managed service, not a better tool | Medium | Kaito needs brand voice, X-TOS-compatible manual posting, compliance escalation — hard for a tool alone | Client quietly replaces 80% of volume with Jasper/ChatGPT |
Verdict: weakly validated (one real paying customer). The open risk is not demand — it's unit economics.
H2 — Risk-assurance demand ("regulated companies will pay a premium for 'every AI output is expert-endorsed'")
| Sub-hypothesis | Confidence | Evidence | What would falsify it |
|---|---|---|---|
| High-stakes domains reject pure AI | High | Industry consensus + regulatory direction | — |
| Companies will hand this work to an external managed service | Low — most dangerous | Zero real customers. Acme Financial is a fictional dogfood; Jumio has no sandbox account | 3 fintech conversations all die on data/liability |
| We can productize "assurance" (legally & commercially) | Unknown | No SLA / indemnity / audit-artifact design exists | Legal review reduces "guarantee" to a disclaimer |
Verdict: unvalidated internal narrative. Directionally supported by the guard-rail face of H1, but it has never faced a real buyer. Engineering investment aligned to H2 (e.g. the KYC tool chain) should be gated on market evidence, not narrative.
Cheapest next tests (in ROI order)
- Kaito unit economics — this week, no code. Expert (Veena) hours/month × loaded cost vs. the monthly fee. One number adjudicates H1's most dangerous sub-hypothesis.
- H2 market probe — mock proposal to 2–3 fintechs. Before building any more KYC plumbing, test whether anyone pays a premium for assurance.
- Kaito renewal/expansion as the H1 strength-of-demand signal.
2. Who We Serve — ICP & Jobs-to-be-Done
Method: synthesized from in-repo evidence — the Kaito engagement corpus (
docs/engagements/), onboarding walkthrough, user stories, dogfood narrative. Findings ranked by evidence strength; "real paid evidence" is explicitly distinguished from "internal narrative."
Findings (evidence-ranked)
F1 — The de-facto ICP: mid-size, high-growth crypto/Web3 companies offloading high-volume, templated, brand-risk-bearing ops work. (Evidence: strong — real payment.) Kaito profile: ~37 employees, ~$35M ARR, growth outrunning headcount. The purchased job: ~30 Launchpad content threads/mo + 48–72h post-launch reply triage (~80% FAQ-shaped).
F2 — The real JTBD is not pure labor arbitrage; it is guard-railed labor arbitrage. (Evidence: strong.) Even in "low-risk" content ops, the engagement is saturated with compliance hard rules: no price/allocation/vesting talk, mandatory refuse-and-escalate on "should I invest," jurisdiction disclaimers, X-TOS-compatible manual posting. The customer isn't buying cheap capacity — they're buying capacity they can trust with their brand in a rule-bound environment. This merges the H1/H2 framing: they are one job's two faces, not two ICPs.
F3 — The customer hires an employee, not a software product. (Evidence: strong.) Delivery is fully anthropomorphic: Kaito's team talks to "Amy" in WhatsApp/Slack; Amy posts via Kaito's own accounts and tools "exactly as a real employee would on Day 1"; in MVP, h.work doesn't even integrate with the distribution channels — the Expert posts manually. What got fired: hiring a junior op (expensive + needs managing), an agency (doesn't know crypto's red lines), ChatGPT + own staff (still consumes own staff; no accountability).
F4 — The risk-assurance ICP (fintech/KYC) is currently pure internal narrative. (Evidence: none.) See H2 above.
F5 — Unit economics is this ICP's hidden red line. (Evidence: adverse signal.) 80% FAQ-shaped volume under always-review burns Expert time on automatable work. The Expert-leverage metric (#3199) shows the team already senses this.
ICP one-pager (what today's evidence supports)
Who: sub-50-person crypto (extendable: fintech) companies whose revenue/growth outruns headcount, with a recurring, monthly-cadence, highly templated external-communication workload that has explicit compliance red lines (campaign content, FAQ triage, KYC inquiries). Trigger: senior staff drowning in this work; unwilling to trust it to agencies or unsupervised AI. They hire h.work to be: "a cheap, reliable digital colleague who knows my industry's red lines — and someone is accountable if it goes wrong." Why they might not buy (untested): price vs. hiring directly; data/liability concerns; volume too small to bother.
3. Competitive Landscape — and Why We Win
The four-layer competitive set, scored against the JTBD
| Dimension | h.work | Pure-AI agents (Sierra / Decagon / Fin) | "AI employee" vendors (Artisan / 11x) | BPO / offshore / VA | Hire directly |
|---|---|---|---|---|---|
| Offload high-volume templated work | Strong | Strong | Adequate (sales-only) | Strong | Weak (expensive) |
| Compliance guardrails / brand safety | Strong (human review + escalation rules) | Weak–Adequate (customer self-configures) | Weak | Weak (QA sampling) | Strong |
| Accountability when wrong | Strong (expert endorsement) | Absent (disclaimers) | Absent | Adequate (contract SLA) | Strong |
| Zero customer-side ops | Strong (fully managed) | Weak (customer builds KB, tunes flows) | Adequate | Adequate (customer trains & supervises) | Weak |
| Industry context (crypto red lines) | Strong (engagement-grade KB) | Weak | Weak | Weak | Adequate |
| Price point | Unknown (anchored to junior salary?) | Decagon median contract ~$400k/yr; Fin per-resolution | Artisan $600–$5,000+/mo | Offshore VA ~$1–3k/mo | Junior op $5–8k/mo + management |
| Gross-margin structure | Weak (always-review burns Expert time) | Strong (pure software) | Strong | Adequate | — |
Pricing evidence (researched 2026-07; goes stale): Sierra sells outcome-based pricing with no public price list and negotiated "outcome" definitions; Decagon is estimated ~$50k/yr platform + ~$0.99/conversation or ~$0.50/resolution, contracts $95k–$590k/yr; Artisan lists $600/mo publicly with enterprise tiers ~$2k–5k+/mo, effectively volume-priced. Sources: Retell, Quiq, OpenNash, eesel, MarketBetter (Artisan), MarketBetter (11x), Landbase.
Key judgments
1. On this JTBD, h.work currently has no head-on competitor. Pure-AI vendors are betting that humans don't belong in the loop — their cost structure and fundraising narrative are locked to de-humanization. Asking Sierra to build an expert network is asking it to attempt suicide-by-pivot — a structural moat, not a feature gap. BPO has humans but no AI leverage and no engineered guardrails. "AI employee" vendors look similar but only do sales outreach, with widely reported delivery-quality disputes.
2. "No head-on competitor" ≠ safe. Threats, ranked:
- ① Customer DIY (ChatGPT + own staff). For AI-native customers this is the default alternative. We win on "doesn't consume your people + someone is accountable." As model capability grows, the arbitrage gap narrows. This is the everyday competition.
- ② Pure-AI vendors moving down-market and bolting on a human-review layer. Structurally hard — but an acquisition of a BPO would shortcut it. Watch their job postings for operations/HITL roles.
- ③ Our own gross margin. If the automation rate never climbs, BPO pricing is our ceiling and pure-AI pricing is our floor — squeezed from both ends. The most dangerous competitor is our own unit economics (closes the loop with H1).
3. Where to differentiate vs. where parity suffices:
- Parity is enough: channel count, chat UI, mobile, notifications — every vendor has these; just don't lose points.
- Must differentiate:
- Productize the assurance — quality receipts, per-reply review evidence, monthly quality reports, escalation SLAs. It's our hardest willingness-to-pay driver and today it's invisible (relates to OD-12).
- Industry red-line engine — Kaito's compliance rules live in Markdown templates today; making guardrails a reusable product asset is the onboarding speed for customer #2.
- Expert-leverage flywheel — capture the Expert edit-diff signal (#3323, ADR-033 L6): automation rate ↑ → margin ↑ → pricing freedom. The single highest-ROI ticket in the company.
The open strategic tension (needs an explicit call)
Always-HITL is simultaneously our moat and our cost structure. Two companies we could become:
- A. "Supervised AI": HITL is transitional; automation climbs; humans mostly exit; high-margin software economics.
- B. "AI-amplified human expertise": humans stay in the loop forever; we sell expert assurance at a premium; never fully automated.
These imply different pricing, ICPs, hiring, and fundraising narratives. Current recommendation (unratified): sell B, build A's engine — market the accountable expert, relentlessly raise automation internally so the flywheel improves our margin rather than eroding our price. This deserves an explicit decision (candidate for OPEN_DECISIONS.md).
A prioritization litmus test
For any feature: "Does it increase Expert leverage, or increase customer trust in the assurance?" If neither — it's probably table stakes; queue it accordingly.
4. Summary — the three questions, answered
| Question | Answer | Confidence |
|---|---|---|
| What problem? Is demand real? | Guard-railed labor arbitrage. Demand weakly validated (n=1 paying); commercial viability hinges on unit economics — unproven. | Medium |
| Who is the customer? Why do they pay? | Sub-50 high-growth crypto/Web3 (fintech unproven); they hire an accountable digital colleague, not software. | Medium-high on crypto ICP; low on fintech |
| Competitive landscape? Why do we win? | No head-on competitor in the "AI leverage + human assurance + fully managed" cell; structural moat = rivals can't/won't add humans. We win by making assurance visible, guardrails reusable, and Experts leveraged. Biggest threats: customer DIY and our own margin. | Medium (pricing data 2026-07) |
Maintenance: re-review at every new signed customer, lost deal, or competitor HITL move. Update evidence labels — this doc's value is its honesty about what is and isn't proven.