Private Essay
Juno Chen · Pricing · PLG · Apr 2026
No password? juno@junochen.com
Six months of pricing decisions — and what they taught me about category, perception, and what happens when there's no benchmark to argue from.
One week before TinyFish's first product launch, I was in front of a whiteboard with our CEO and CPO trying to explain a competitive analysis that didn't add up. Open-source browser automation: free to self-host, per step on cloud. Research APIs: $5 to $2,400 per thousand requests, tiered by complexity. An agent framework with no price at all. Browser infrastructure: by the hour. Same category on every pitch deck. Four completely different answers to the same underlying question. What I'd built wasn't a benchmark. It was a record of four companies making different bets about what an AI agent was worth.
There were three of us in the room. No pricing committee, no cross-functional process. Every call we'd make for the next six months would be made the same way: me, the CEO, the CPO, the data we had, and a judgment call. Our CEO said the only useful thing anyone said that week: "No one knows what's the best way for pricing yet." I heard it as defeat. It was permission — and a design brief.
SaaS pricing has twenty years of receipts. AI agent pricing has none. The ambivalence you feel opening a competitor's pricing page isn't a gap in your knowledge. It's the actual condition of the market. Three months into my role, we launched with three bets. The three months after taught me three more.
Pricing is a structural decision baked into the first cohort. The unit you choose shapes who self-selects. Who self-selects shapes the data you collect. Get it wrong and you're not iterating toward a better answer — you're collecting noise from the wrong customers and calling it signal.
I pushed for the cost work before we built the pricing page. Not because I had time — I had a week. Because any number we put on screen without this was a guess, not a price.
There was no moment the cost model was done. Every infrastructure decision, model upgrade, and new task type shifted the curve. I stopped treating it as a deliverable and started treating it as production infrastructure — the way an engineer maintains a system, not the way an analyst delivers a report.
The cost model isn't the prerequisite to pricing. It is pricing — running continuously in the background.
Going slow cost us a delayed launch. Going fast would have cost us something worse: pricing from unvalidated assumptions, then discovering the gap at scale. What users eventually see and trust — the credit display, the cost estimate, the charge confirmation — sits on top of this cost layer. You can't design the visible surface honestly without understanding the invisible one first.
Flat task pricing hides a subsidy. Simple workflows run fast and cheap; complex ones run long and expensive. Charge the same rate per task and simple-workflow customers unknowingly cover the cost of complex-workflow customers. Eventually they notice. Then they leave.
Agent steps removed the subsidy. One action, one step. Tracking a package: 2 steps. Booking a flight: 19. A complex multi-field form: 35. Cost follows complexity. Price follows cost.
What I didn't anticipate: a unit isn't only a billing decision. It determines which competitive room you walk into before you write a word of copy. Per-GB puts you next to web scrapers. Per-token puts you alongside developer infrastructure. Per-step puts you in the AI agent category. We weren't running a positioning exercise — we were building a pricing page. The question was identical: which room do we want to be evaluated in?
The unit answered it before any copy did. Then we had to furnish the room. Choosing a rare unit also meant choosing to educate. The gap between a unit that's honest and one that's legible is a product problem. I underestimated it at the time and built toward it later.
"AI introduces variable costs that can fluctuate widely depending on the task or model used. Our approach is to separate those elements — we pass through the inference cost transparently to the customer and layer our pricing around the value we deliver. That way, our incentives stay aligned: we're not optimizing for cheaper models, but for better outcomes."
— Jasdeep Garcha, Vercel
The market had already run the experiment we didn't want to run. Credit-based billing with unpredictable consumption creates a specific failure: users who can't estimate what a task will cost before they start stop starting. They don't cancel. They freeze.
When you can't predict what a task will cost before you start, you stop starting. The billing model breaks before the product does.
"Instead of 1,000 credits, they deducted 9,000."
— User review, competitor platform
Failed runs cost zero credits. The design decision was about clarity before it was about fairness.
What I hadn't anticipated: what outcome billing would do inside the company. The moment failed runs were free, task success rate became everyone's metric — without anyone announcing it. Engineering, product, customer success: all with direct financial skin in the game on whether the agent actually finished the work. The 30-day task success rate moved from an engineering dashboard into every weekly review.
The incentive ran in both directions. On the customer side: free failure meant low stakes to start. Users ran more tasks, explored more edge cases, discovered what the agent was actually good at — instead of stopping after one billing shock and leaving. The structure created the behaviors we wanted on both sides of the product.
Because we launched without feature walls — no gated capabilities, no artificial limits — customers could show us what they actually valued rather than tell us. Two patterns emerged that I hadn't modeled. An enterprise buyer: "I don't need one agent — I need fifty running simultaneously, or this doesn't fit my workflow." Concurrency wasn't a feature request. It was the definition of value for that segment — the difference between a tool and a workflow. The second: every power user who upgraded had already succeeded at scale. The tier just named what they'd proven. Neither lever was designed. Both were discovered because the product got out of the way.
This is what commitment-by-design means in practice: the promise you build into the product structure shapes what the organization optimizes for. We hadn't designed a pricing policy. We'd designed a commitment — and a commitment has consequences on both sides.
I built the billing infrastructure before the visibility layer. I could track exactly what any customer owed. I couldn't show them — before the charge — what they were spending and why.
Customers received charges they couldn't independently verify. Even when the number was small, the reaction was consistent: "This feels expensive." When I shipped the real-time usage display, the unit explainer, and a cost calculator, the reaction inverted. Same product. Same price. Same unit.
"Wait — that's actually it?"
Nothing changed except what they could see. The shift was entirely about what they could verify.
In agentic systems, the primary UX challenge is making the invisible visible — what the agent did, why it cost what it cost, whether the output holds up. The billing infrastructure and the visibility layer were never separate problems. I just sequenced them that way.
"We emphasize real-time transparency, giving them a
— Annika Schmid, PostHog/usagecommand to see the exact cost of their conversation. Every customer can set a hard billing limit — guaranteeing they never go over the budget they set."
Then customers asked for PAYG — the right call. But PAYG makes transparency more critical, not less. A monthly plan has a ceiling. PAYG doesn't. Without real-time visibility, PAYG creates exactly the billing shock I'd watched accumulate in one-star reviews, through a different mechanism. Transparency and PAYG aren't two features. They're the same obligation.
PAYG also wasn't a pricing decision — it was a product decision. Real-time metering meant migrating our billing stack to our own infrastructure. A line item became an engineering sprint. The cost wasn't just the sprint. It was the months of demand we couldn't capture while the infrastructure wasn't ready. It showed in retention numbers before it showed anywhere else.
Pricing decisions have product consequences. The earlier you make them, the cheaper those consequences are.
Most sophisticated programs go through the same arc. They launch with descriptive names — names that try to explain what you get. Then they simplify toward resonance: names that signal who you are and who belongs.
I've seen this across loyalty programs, premium memberships, and enterprise SaaS. The most common mistake isn't a bad product — it's a name that recruits the wrong person. A tier called "Professional" attracts people who self-identify as professionals. A tier called "Developer" attracts developers. Get that match wrong and you're not just losing conversions — you're onboarding someone into a product that will disappoint them, then blaming the product for the mismatch the name created.
The name isn't a description. It's an invitation. Invitations are precise — they tell the right people they belong, and the wrong people they don't. The tier card is the first copy a customer reads. Get the words wrong and no amount of onboarding corrects the mental model that forms in the first three seconds.
We learned this when paid users were underperforming free users on the same tasks. The tier name had attracted users with completely different expectations. We renamed toward outcome language — what users would have, what was predictably theirs. The product didn't change. The right customers finally found it.
One label, one wrong mental model. No amount of onboarding fixes what the name broke before anyone logged in.
As the product expanded to multiple surfaces, the instinct was to price each precisely — a different unit for each product. Defensible. Theoretically correct. The case for it was real.
The problem: every new pricing surface is a new mental model the customer has to learn. My contribution was working out the translation layer — a unified credit unit across surfaces, with conversion ratios underneath. The customer always sees one thing. The ratios vary; that's my problem to maintain, not theirs to learn.
This is what a design token does: it presents a unified surface while letting the implementation vary underneath. The abstraction protects users from the internal architecture. Every new surface that speaks the same language is one more reason a user's mental model doesn't break — one more reason they stay rather than re-evaluate the relationship.
The trade is real. I maintain a conversion table that grows with every new product — one of the less glamorous artifacts of the decision, and the most durable.
Precision that doesn't survive the next product you ship isn't precision. It's technical debt in the pricing model — and in the design system.
The bets you make before you ship are shaped by the data you have. The corrections you make after are shaped by the data you didn't know you'd need. Both matter. The difference is knowing which one you're doing at any given moment.
The question never really changed. Every bet was a version of the same one: what does the person on the other side need to trust what they can't see? I've been answering that question for fifteen years — in supply chains, in B2B marketplaces, in AI agent interfaces. The pricing page taught me it doesn't only live in the interface. It lives everywhere the product makes a promise.