Showing posts with label AI Agents. Show all posts
Showing posts with label AI Agents. Show all posts

Friday, 17 April 2026

Stop Reaching for Agents

Every week I see another team announce they're "building an agent" for a problem that a single well-written prompt would solve. A few weeks later, they're debugging a loop where the model keeps calling the wrong tool, blowing through tokens, and producing answers worse than the one-shot baseline they skipped past.

This is the default failure mode of LLM engineering right now. The industry keeps pushing toward the flashiest pattern on the menu, and teams keep mistaking complexity for capability. The truth is boringly simple: the right pattern is almost always the simplest one that works, and you should have to be forced up the ladder, not invited.


Framework you can use 

Think of LLM patterns as rungs on a ladder. Each rung adds capability, but also adds cost, latency, failure modes, and debugging surface area. You climb only when the rung below genuinely can't do the job.



Rung 1 — Single prompt. Zero-shot or few-shot. One call, one answer. This is your starting point for every task, without exception. Modern frontier models are astonishingly capable in a single call, and most teams underestimate how far good prompting alone can take them. 

Examples: classifying emails as urgent/normal/spam, drafting a reply to a customer message, summarizing a meeting transcript into action items, extracting fields from a contract into JSON.





Rung 2 — Structured prompting and chain-of-thought. When the model gets answers wrong because it's skipping reasoning steps or producing messy output, you don't need a new architecture. You need better instructions. Ask it to think step by step, give it a structure to fill in, show it examples of the reasoning you want. This fixes more problems than people expect. 

Examples: math word problems where the model jumps to a wrong answer, multi-criteria decisions like "should we approve this expense" where you want the reasoning shown, data extraction tasks where output format matters.





Rung 3 — Retrieval-augmented generation (RAG). When the model doesn't know something — your internal docs, fresh data, domain-specific knowledge — bolt on retrieval. You're not changing how the model thinks, just what it has access to. RAG is often mistakenly treated as the default for any knowledge-heavy task; it's the default only when the knowledge genuinely isn't in the weights.

  Examples: answering questions from your company's internal wiki, a legal research tool grounded in a specific case database, a coding assistant that needs to reference your private API documentation, a support bot that cites current policy docs.








Rung 4 — Workflows. Prompt chaining, routing, parallelization. You use these when the task has distinct sub-tasks that you can enumerate in advance. Classify the input, then draft, then check. Or: run these three analyses in parallel and synthesize. The defining feature of a workflow is that you wrote down the steps. The model fills in each one, but the control flow is yours. 

Examples: a translation pipeline that translates, then checks for cultural appropriateness, then adjusts tone. A customer inquiry system that first routes the message to sales/support/billing, then dispatches to a handler tuned for that category. A document analyzer that extracts entities, sentiment, and topics in parallel, then synthesizes a report. A content moderation flow where a draft is generated, then evaluated against policy, then revised if flagged.















Rung 5 — Agents. An LLM in a loop with tools, deciding what to do next. You use this when the path genuinely isn't knowable in advance — the model has to observe, decide, act, observe again. Agents are powerful and they're the right answer for some problems, but they're expensive, slow, and the hardest pattern to debug. If you can write down the steps, you don't need an agent; you need a workflow.

 Examples: a coding assistant that explores an unfamiliar codebase to fix a bug, where the next file to open depends on what it just read. An open-ended research task where findings from one search determine the next query. A browser agent completing a multi-step booking where page contents dictate the next click. Incident response where the diagnostic path branches based on what each check reveals.




Rung 6 — Fine-tuning. Last resort. Use it when prompting has plateaued, you have a stable task, and you have real data. Fine-tuning trades flexibility for performance on a narrow distribution, and the maintenance cost is real. Most teams who think they need fine-tuning actually need better prompts or better retrieval. 

Examples: matching a very specific brand voice across millions of generated product descriptions, a narrow classification task with labeled data where prompting plateaus below required accuracy, replicating a structured output format that few-shot examples can't reliably produce.


Decision questions

Instead of picking a pattern, ask these questions in order and let them pick for you:

Does the model know enough? If the task requires knowledge the model doesn't have — private documents, today's data, niche domain details — you need RAG. If it has the knowledge, skip this rung.

Can one prompt do it well? Try it before assuming it can't. You'd be surprised how often "I need a multi-step pipeline" turns into "actually, one prompt with good structure handles it." If a single prompt works, ship it.

Can I write down the steps in advance? This is the workflow-vs-agent line, and it's the most important question in the whole framework. If you can enumerate the steps — even if there are branches — you want a workflow. Hardcode the control flow, let the model handle each step. You get deterministic behavior, easier debugging, lower cost, and predictable latency. 

Agents are for when the steps genuinely can't be known ahead of time.






Do the steps depend on each other? Sequential steps become a prompt chain. Independent steps run in parallel. Steps that depend on the input type get a router at the front.

Is quality inconsistent? Add an evaluator-optimizer loop — one model generates, another critiques, the first revises. This is often the right fix before reaching for anything more complex.



Have I plateaued on everything else? Only then does fine-tuning enter the conversation.


Why the simple-first bias matters

There are three practical reasons the ladder approach beats jumping straight to complex patterns, and they compound.




The first is cost. Every additional LLM call, every tool invocation, every agent loop iteration multiplies your token spend. A workflow with three sequential calls costs 3x a single prompt. An agent that takes ten loops to converge costs 10x — and that's when it converges. In production, cost differences of 10–100x between patterns are common.

The second is reliability. Every LLM call has some failure rate. Chain five calls together and you compound those failures. Agents, which can loop arbitrarily, compound them worst of all. Simpler patterns have fewer places to fail and fewer places where a failure cascades.

The third is debuggability. When a single-prompt system gives a bad answer, you change the prompt. When an agent gives a bad answer, you stare at a 40-step trace trying to figure out which decision went sideways, whether the tool returned the wrong thing, whether the model misread the tool output, whether the loop should have terminated earlier. The complexity you added to solve the problem becomes the problem.



Worked examples

Customer support from product docs. A team reaching for the hot pattern might build an agent: it plans a query strategy, searches docs, reads pages, decides whether to search again, drafts an answer, self-critiques, and revises. Lots of moving parts. Impressive demo.

The ladder approach asks the questions instead. Does the model know your docs? No — so you need retrieval. Can one prompt do it well once the docs are retrieved? Usually yes: "Here's the user's question, here are the relevant doc passages, answer using only the passages." Can you write down the steps? Yes: retrieve, then answer. That's a two-step chain. No agent, no loop, no self-critique — unless measurement shows you actually need them.

Nine times out of ten, the two-step chain ships faster, costs a fraction as much, is easier to debug, and performs as well or better than the agent. 

The tenth case — where questions are genuinely open-ended and require multi-hop reasoning across documents — is where an agent might earn its keep. But you discover that by measuring, not by assuming.

Generating weekly sales reports. Someone pitches an agent that gathers data, analyzes it, and writes the narrative. But walk through the questions. Does the model know your sales data? No — but you don't need RAG either; you need a direct query to your database. Can one prompt do it well? Almost: given the raw numbers, a single prompt can produce a decent narrative. Can you write down the steps? Completely: pull the data, format it, ask the model to write the narrative, optionally ask a second call to check the numbers match. That's a fixed workflow, not an agent. You know exactly what happens every Monday at 9am.

Debugging a failing test in an unfamiliar codebase. Now the agent is justified. Does the model know the codebase? No. Can one prompt do it well? No — the model needs to look at actual code. Can you write down the steps? This is where it breaks down. The next file to open depends on what the last file contained. The error might be in the test, the code under test, a shared dependency, or a config file. You can't enumerate the path because the path depends on what's found along the way. This is the shape of a problem that actually needs an agent: genuine dynamic exploration, not a pipeline dressed up in a loop.


The habit to build

When you pick up a new LLM task, resist the impulse to architect. Start at the bottom of the ladder. Write the simplest prompt that could plausibly work, run it on real examples, and see what breaks. Let the failures tell you which rung to climb to. The specific failure mode — "it doesn't know our product," "it skips reasoning steps," "it can't decide which analysis to run" — maps cleanly onto the next rung.

This is a less glamorous way to build, but it's how you end up with systems that actually work in production. The goal isn't to use the most sophisticated pattern. The goal is to solve the problem with as little machinery as possible, because every piece of machinery is something that can go wrong at 3am.

Start simple. Climb only when you're forced to Ship.

Friday, 10 April 2026

Agents Don't Speak It

This is my continuation post of Broken promise of Agile and Agile Manifesto In Age Of Ai Agentic Software development world

What happens to sprint planning, standups, retros, and bug reports when the team building your software isn't human.

Jeff Bezos had a simple heuristic for team size: if two pizzas can't feed the team, the team is too big. It was never really about pizza. It was about communication overhead — the invisible tax that grows quadratically as you add people. Small teams move fast because coordination is cheap.

Now imagine replacing those six engineers with six AI agents. No standups. No slack threads at midnight. No pushback during planning. They just run week day / weekend / 24*7 .

Sounds like a superpower. It isn't — or rather, it isn't straightforwardly one. The coordination problems don't disappear. They move , they concentrate and they become invisible in ways that human teams never were.




Fundamental Difference

When you manage six engineers, you get a huge amount of coordination intelligence for free.


Btw what is coordination intelligence ? 

Coordination intelligence is the ability to self-organize around incomplete information without being told to — noticing collisions, resolving ambiguity, pushing back before work goes wrong. In human teams it emerges for free from social context: reputation, embarrassment, shared history


 Engineers notice when two people are working on the same thing. They push back on bad estimates. They carry context from last month's decision. They feel embarrassed when they ship something broken. That embarrassment is load-bearing infrastructure.

Agents have none of this. An agent will accept any scope you give it, work confidently in the wrong direction for hours, produce six internally-consistent but mutually-incompatible outputs — and report back with no signal that anything went wrong.

Six agents don't reduce management overhead. They concentrate it into a single engineer's head.


What Happens to Each Ceremony

Sprint Planning

Planning with engineers is a negotiation. Engineers push back. That pushback is annoying — it's also your earliest warning system. Agents don't negotiate. They accept any scope. Without pushback you'll consistently over-assign, and agents won't tell you — they'll just produce something confidently wrong at scale.

Sprint planning stops being about capacity negotiation. It becomes context package design. 


Lets expand Context package 

Context package design is the discipline of deciding exactly what an agent needs to know, what it must not know, and where its work begins and ends — so it can complete a task correctly without asking questions, without drifting into adjacent scope, and without conflicting with what other agents are building in parallel.


For each task: what does this agent need to know? What must it explicitly not know? Where does it hand off, and to whom? The role shifts from breaking down stories to writing intelligence mission briefs.

Daily Standup

Standups exist to catch invisible blockers early through human signal — tone, hesitation, the "I'll figure it out" when someone won't. Agents don't have tone. The standup equivalent becomes a health-check dashboard: are agents producing output? Did any contradict each other? Is any stuck in a tool-call loop? Status collection becomes anomaly detection.

Sprint Review

In a normal review, the engineer who built the feature explains its edge cases. Knowledge transfers. Pride is a quality signal. With agents, the output exists but nobody fully understands it. An agent can produce 600 lines of passing code and the engineer who prompted it cannot explain every architectural decision.


CeremonyHuman purposeWith agents, becomes
Sprint PlanningCapacity negotiation + pushbackContext package design — precise briefs, explicit scope boundariesmutates
Daily StandupCatching invisible blockers via tone + signalAnomaly detection dashboard — traces, diffs, loop detectionmutates
Sprint ReviewDemo + informal knowledge transferComprehension gate — a human must own and explain the outputintensifies
RetrospectiveProcessing human failures via memoryPrompt autopsy — full trace replay, brief quality analysismutates
Bug TriageAssign → investigate → fix with moral ownershipRe-ownership ritual before fix — someone must read the whole moduleintensifies
Code ReviewPeer knowledge transfer + quality gateReview wall — agents outpace human review capacity almost immediatelybreaks

Human tech debt traces back to a decision. Agent tech debt has no author intent — only output.


Bug Reports

Bug reported → assigned to whom? The agent that wrote the buggy code no longer exists. The bug might live in the interaction between two agents' outputs — nobody's fault in isolation. If you assign the fix to another agent without human comprehension in between, you risk entering a patch spiral.




Code Review

Agents generate PRs faster than a single human can review them. The review wall hits almost immediately. What Options you have ? 



New Pattern Emerging

Every Agile ceremony was designed to solve a human coordination problem. When you replace engineers with agents, those problems don't disappear — they move up the stack to the one or two humans managing the agents. Those humans now carry the full cognitive load that was previously distributed across a team of six.



The skill of managing agents isn't delegation. It's context architecture — what each agent knows, when, and in what form.


Agile solved for human limits: attention, memory, communication bandwidth. Agents don't have those limits. But they have different ones — context window coherence, statelessness, silent failure, no social accountability. We don't yet have a name for the ceremonies that solve for those.

The teams who figure out the new paradigm first will ship faster — not because they have more agents, but because they've rebuilt the coordination layer from scratch for the law of physics that actually govern them.


Thursday, 8 January 2026

The Circle Game: Why Everyone's Wrong About AI's Circular Finance

Financial press is filled with panic attack about "circular finance" in AI. Nvidia invests in OpenAI, OpenAI buys Nvidia chips. Microsoft funds Anthropic, Anthropic rents Microsoft Azure. Oracle backs AI labs, labs fill Oracle datacenters.

"This is a bubble!" they cry. "Circular financing!" they warn. "Just like the dot-com crash!"

But here's what nobody is telling that : This is literally how banking has worked for centuries.

And banks are rewarded for it.



Let me explain why the "circular finance" criticism reveals more about financial illiteracy than about AI sustainability.

How Banking Actually Works: The Original Circle

Let's start with what everyone accepts as normal, healthy finance:

Traditional Business Loan:

  1. Bank lends you $500,000 to start a restaurant
  2. You use that $500,000 to buy equipment, inventory, lease space
  3. You operate the restaurant and generate revenue
  4. You pay the bank back $600,000 over 5 years (principal + interest)
  5. The bank's balance sheet grows by $100,000

Wait, isn't this circular?

Bank gives you money. You spend that money. You make money from what you bought. You give money back to the bank. Same money just moved in a circle, but the bank's numbers went up.

Nobody calls this a "circular finance scheme" or worries it's unsustainable. Why? Because we understand the mechanism:

  • Bank provided capital
  • Capital bought productive assets (kitchen equipment, inventory, house, education loan, corporate loan)
  • Productive assets generated revenue
  • Revenue exceeded costs
  • Bank captured a portion of that value creation (interest)

Productive finance. Money circulates, but value is created in the process. The bank's growing balance sheet reflects real economic activity, not financial engineering.

How AI Finance Actually Works: The Same Circle

Now let's look at what everyone's panicking about:

AI Lab Financing:

  1. Nvidia invests/lends $50B to an AI lab
  2. AI lab uses $50B to buy Nvidia chips and datacenter capacity
  3. AI lab builds models and generates revenue from AI services
  4. AI lab pays Nvidia back through revenue, equity appreciation, or future purchases
  5. Nvidia's balance sheet grows

This is the exact same structure as banking.

Replace "bank" with "Nvidia" and "restaurant equipment" with "AI chips" and you have the identical circular flow:

  • Nvidia provides capital
  • Capital buys productive assets (GPUs, compute infrastructure)
  • Productive assets generate revenue (AI services, API calls, subscriptions)
  • Revenue exceeds costs (hopefully)
  • Nvidia captures a portion of that value creation (returns, revenue)

Money circulates. Nvidia's numbers go up. But if the AI services generate real revenue from real customers paying real money, this is productive finance—not a house of cards.

Why the Circle Works (When It Works)

Both banking circular finance and AI circular finance succeed under the same conditions:

1. Productive Asset Purchase

Banking: Loan buys equipment that produces valuable goods/services 

AI: Investment buys chips that produce valuable AI capabilities

2. Real Customer Demand

Banking: Customers pay for restaurant meals, manufactured goods, services 

AI: Customers pay for AI capabilities, productivity gains, automation

3. Revenue > Costs

Banking: Business generates enough revenue to cover operations + loan repayment 

AI: AI lab generates enough revenue to cover compute costs + investor returns

4. Risk-Adjusted Returns

Banking: Bank charges interest rate that compensates for default risk 

AI: Nvidia prices investments/chips to compensate for business risk

When these conditions hold, the circle is sustainable. When they don't, it collapses—whether you're lending to restaurants or AI labs.

Real Question Isn't "Is It Circular?" But "Is It Productive?"

Entire circular finance critique misses the fundamental question: Are the acquired assets generating real value?

Bad circular finance (dot-com era):

  • Lucent lent money to telecom companies
  • Telecom companies bought Lucent equipment
  • Equipment sat unused because demand was overestimated
  • Telecoms couldn't generate revenue to pay back loans
  • Lucent wrote off billions, 47 carriers went bankrupt
  • The circle broke because the assets weren't productive

Good circular finance (banking forever):

  • Bank lends money to businesses
  • Businesses buy productive equipment
  • Equipment generates goods/services customers want
  • Revenue pays back loans
  • Everyone prospers
  • The circle continues because the assets ARE productive

AI circular finance (TBD):

  • Nvidia/Oracle/Microsoft fund AI labs
  • Labs buy compute infrastructure
  • Infrastructure produces AI capabilities
  • Customers pay for those capabilities (or don't)
  • Revenue justifies infrastructure costs (or doesn't)
  • The circle continues IF the assets are productive

What the Numbers Actually Show

Let's look at what OpenAI's commitments actually mean:

The "Scary" Numbers:

  • OpenAI revenue projection 2025: $13B
  • Infrastructure commitments: $300B (Oracle) + $90B (AMD) + $38B (AWS) = $428B
  • Ratio: 33x annual revenue in commitments

But here's the context nobody mentions:

These are multi-year commitments, not immediate spending. If spread over 10 years, that's $42.8B/year. If OpenAI grows revenue at 50% annually (slower than recent growth), they hit $118B annual revenue by year 5.

Compare to restaurant financing:

  • Restaurant revenue Year 1: $500k
  • Bank loan: $500k (1x annual revenue)
  • Equipment lifespan: 10 years
  • If restaurant grows 20%/year, loan becomes 0.16x revenue by year 10

The structure is identical. The only question is: Will AI services revenue grow fast enough to justify infrastructure investment?

That's not a question about circular finance. That's a question about market demand and business fundamentals.

Why Nvidia Isn't "Playing Bank" Wrong

Critics say Nvidia is taking excessive risk by being both investor and supplier. But this is actually standard practice in capital-intensive industries:

Equipment Financing Examples:

  • John Deere Financial lends money to farmers to buy John Deere tractors
  • Caterpillar Financial lends money to construction companies to buy Caterpillar equipment
  • Boeing Capital lends money to airlines to buy Boeing planes
  • Tesla provides financing for Tesla solar installations

In every case:

  1. Manufacturer has capital
  2. Manufacturer lends to customers
  3. Customers buy manufacturer's products
  4. Products generate revenue for customers
  5. Customers pay back manufacturer
  6. Manufacturer's business grows

This is called vendor financing, and it's normal.

Nvidia providing capital for customers to buy Nvidia chips is the AI equivalent of John Deere financing farmers to buy tractors. The circle works if tractors generate farming revenue and chips generate AI service revenue.

The Real Risks (Which Have Nothing to Do With Circularity)

I'm not saying AI finance is risk-free. The real risks are:

1. Demand Risk

Will enterprise customers pay enough for AI services to justify infrastructure costs? This is the same risk restaurants face—will customers pay enough for meals to justify kitchen equipment costs?

2. Competition Risk

Will AI service margins get compressed by competition? This is the same risk any business faces when competitors enter the market.

3. Technology Risk

Will current infrastructure become obsolete before generating sufficient returns? This is the same risk farmers face when buying tractors that might be superseded by better equipment.

4. Execution Risk

Will AI labs successfully build profitable businesses? This is the same risk any loan carries—will the borrower execute their business plan?

None of these risks are about circular finance. They're standard business risks that exist in any capital-intensive industry.

Why the Criticism Persists

If AI circular finance is structurally identical to banking and vendor financing, why is everyone panicking?

Three reasons:

1. Scale Shock

The numbers are enormous. $428B in commitments sounds scary. But in context, it's not crazy:

  • Global banking assets: $180 trillion
  • Amazon's capital investments 2019-2023: $240B
  • Meta's planned datacenter spend: Similar scale
  • AI infrastructure is capital-intensive, like all infrastructure

2. Speed Shock

AI scaled incredibly fast. OpenAI went from research lab to $13B revenue in ~5 years. Traditional businesses take decades to reach this scale, so the financing mechanisms feel rushed.

3. Misunderstanding "Circular"

People see money flowing in a circle and assume it's artificial. But ALL productive finance is circular—money flows from capital providers to businesses to customers back to capital providers. That's called an economy.

The Actual Test

Whether AI circular finance succeeds or fails will depend on one question:

Do AI services generate enough customer value to justify the infrastructure investment?

If yes:

  • Labs make revenue from real customers
  • Revenue pays infrastructure providers
  • Infrastructure providers profit
  • Circle continues sustainably
  • Just like banking for centuries

If no:

  • Labs can't generate sufficient revenue
  • Infrastructure providers don't get paid back
  • Investments written off
  • Circle breaks
  • Just like failed business loans

This isn't novel financial engineering. It's not a bubble indicator. It's not concerning because it's circular.

It's just finance. The same finance that's funded every capital-intensive industry from railroads to restaurants to rental cars.

What This Really Reveals

The "circular finance" panic reveals something uncomfortable: Most people don't understand how productive finance works.

They see money flowing in circles and panic because they think money should flow in straight lines. But productive economies have ALWAYS been circular:

  • Banks lend to businesses
  • Businesses buy equipment
  • Equipment generates goods
  • Customers buy goods
  • Revenue pays back banks
  • Banks lend more

The circle is a feature, not a bug. It's called economic growth.

The question isn't "Is AI finance circular?" (Yes, like all finance)

The question is "Is AI infrastructure productive?" (Are customers willing to pay for AI services?)

If you're worried about AI circular finance, ask yourself: Are you also worried about every business loan a bank makes? Because structurally, they're the same thing.

If the answer is no, then your concern isn't about circular finance. It's about whether you believe AI services will generate sufficient customer demand.

That's a valid concern. But let's be honest about what we're actually debating.

The Bottom Line

Banks have done circular finance for centuries and we reward them for it. They lend money, borrowers buy productive assets, assets generate revenue, revenue pays back loans, banks profit.

Nvidia is doing the same thing with AI labs. They provide capital, labs buy chips, chips generate AI services, services generate revenue, revenue flows back to Nvidia.

The structure is identical. The mechanism is identical. The only question is execution: Will AI services generate enough demand?

If yes, this is productive finance that creates value. If no, it's failed investment that destroys value.

But calling it "circular finance" as if that's inherently problematic just reveals you don't understand how banking—or business—has ever worked.

The circle isn't the problem. The circle is how capitalism functions.



Thursday, 13 November 2025

Agentic Commerce War - Car with bumpy roads

 There's a lawsuit that should make every engineer building AI applications pause and think carefully about the world they're creating. Amazon is suing Perplexity AI, and while the legal complaint talks about "covert access" and "computer fraud," what's really happening is far more interesting: we're watching the first shots fired in a war over who gets to control the future of commerce.

"AI revolution" in shopping is probably going to make markets less competitive, not more. Let me explain why.




The Pattern-Matching Disguised as Innovation



We've seen this movie before. In the 2000s, it was about who controls app distribution. Apple and Google built "open" platforms, welcomed developers, then extracted 30% rent from everyone. In the 2010s, it was about who controls attention. Facebook and Google became the gatekeepers to your customers, then jacked up ad prices once you were dependent on them.

Now we're in the 2020s, and the game is about who controls shopping intent. The technology changed—from apps to ads to AI agents—but the fundamental power dynamics remain depressingly familiar.

Agentic commerce is the fancy term for AI systems that can shop on your behalf. Tell ChatGPT you need running shoes under $100, and it searches stores, compares options, and completes the purchase. No browsing, no clicking through pages, no "adding to cart." The AI does it all.

McKinsey forecasts this could generate $1 trillion in global commerce by 2030. Traffic to U.S. retail sites from GenAI browsers already jumped multi fold year-over-year in 2025. 

Amazon vs. Perplexity: A Case Study in Platform Power

Here's what actually happened, stripped of the legal jargon:

November 2024: Amazon catches Perplexity using AI agents to make purchases through Amazon accounts. They tell Perplexity to stop. Perplexity agrees.

July 2025: Perplexity launches "Comet," their AI browser that can shop for you. Price tag: $200/month.

August 2025: Amazon detects Comet's agents are back, but this time they're disguised as Google Chrome browsers. Amazon implements security measures to block them.

Within 24 hours: Perplexity releases an update that evades Amazon's blocks.

November 2025: Amazon files a federal lawsuit accusing Perplexity of violating the Computer Fraud and Abuse Act. Perplexity publishes a blog post titled "Bullying is Not Innovation."

Now, you might think this is about security or customer protection. And sure, those are real concerns—when AI agents access customer accounts, make purchases, and handle payment data, security matters enormously.

But let's be honest about what's actually happening here: Amazon is defending its moat.




Amazon built a trillion-dollar business by owning the customer relationship. They know what you buy, when you buy it, how much you're willing to pay, and what you'll probably want next. This data advantage is what makes Amazon Rufus (their own shopping agent) dangerous to competitors—it already knows you better than any third-party agent ever could.

If Perplexity's agents can freely roam Amazon's platform, comparison-shop ruthlessly, and complete purchases without Amazon controlling the experience, then Amazon loses three critical things:

  1. The ability to show you ads for products they want you to buy
  2. The ability to promote their own private-label brands
  3. The data about what AI-assisted shopping actually looks like

This is Amazon's "app store moment." And they learned from Apple: if you're going to allow third parties to build on your platform, you need to control who gets access and extract rent from those you approve.

Architecture of Control: How This Actually Works

Let's talk about the technical stack for a moment, because this is where it gets interesting from an engineering perspective.

The Five-Layer Problem

Layer 1: Consumers delegate shopping tasks to AI agents, often paying $20-200/month for the privilege.

Layer 2: AI Agents (ChatGPT Operator, Perplexity Comet, Amazon Rufus) search, compare, and transact on your behalf.

Layer 3: Trust & Payment Infrastructure (Visa, Mastercard, Stripe) verify agent identity and process payments.

Layer 4: Platform Gatekeepers (Amazon, Google, Apple) control access to inventory and customer data.

Layer 5: Merchants & Brands fulfill orders and watch their margins compress.

Where power concentrates: not at the AI layer where everyone's focused, but at Layer 3 (payments) and Layer 4(Platform Gatekeepers)




Why Payments Matter More Than You Think

Visa and Mastercard are quietly positioning themselves as the critical trust infrastructure for agentic commerce. They're partnering with Cloudflare to implement Web Bot Auth—a cryptographic authentication protocol that lets merchants verify which AI agents are legitimate.

Think about the implications: if every agentic transaction must flow through payment network authentication, then Visa and Mastercard become the gatekeepers of which agents can transact at all. They've turned themselves into the identity verification layer for AI agents, which means they can collect tolls on the entire ecosystem.

This is brilliant infrastructure play. While everyone's fighting over the AI layer, the payment networks are becoming the new platform.

The Security Nightmare Nobody Wants to Talk About

Here's the thing that keeps security engineers up at night: traditional fraud detection assumes humans are making purchases. You can look at behavioral patterns, device fingerprinting, velocity checks—all the usual signals that distinguish legitimate users from attackers.

But what happens when the "legitimate" user is an AI agent that behaves like a bot because it is a bot?

The attack surface is enormous:

  • Agent manipulation: Increase vulnerability rate for tricking AI agents with fake listings or manipulated reviews
  • Automated account takeover: AI can run credential stuffing attacks at scale, then use compromised accounts to make "legitimate" agent purchases
  • Synthetic identity fraud: Generate deepfakes and fake identities that pass agent verification
  • Phishing at industrial scale: AI-generated personalized phishing that tricks both humans and other agents

To successfully implement agentic commerce, you need to solve the impossible problem: identify agents, distinguish legitimate from malicious ones, verify consumer intent, and do all of this in real-time at massive scale.

This isn't just a "hard problem"—it requires fundamentally rethinking identity, authentication, and trust in ways our current infrastructure wasn't designed for.

Legal Black Hole

The most fascinating aspect of this entire situation is that nobody knows what the law actually says about AI agents making purchases on your behalf.

Consider this scenario: Your AI agent buys you running shoes. They don't fit. Who's responsible?

  • Is the AI agent your "employee" acting under your authority?
  • Is it a contractor working for the AI company?
  • Did you actually "agree" to the purchase, or did the AI misinterpret your intent?
  • Can you return them under standard return policies, or do different rules apply?

The Uniform Electronic Transactions Act (UETA) and E-SIGN Act validate electronic signatures and contracts, but they were written assuming humans click "I agree." They don't tell us how to handle situations where an AI system makes autonomous decisions based on high-level instructions like "buy me running shoes under $100."

And it gets worse. When things go wrong—the agent buys the wrong product, accesses the wrong account, or exposes payment data—who's liable?

The legal frameworks assume someone clicked a button and agreed to terms. But with agentic AI:

  • The consumer gave high-level intent ("I need shoes")
  • The AI developer built the agent with certain objectives
  • The platform (Amazon) sets rules about what's allowed
  • The payment processor enables the transaction

When something breaks, you've got four parties pointing at each other saying "not my fault."

This isn't edge case stuff—this is the fundamental contract law question that needs answering before any of this scales. And right now? It's a complete void.

 Three Scenarios 

Based on the current trajectory, few things could happen:

Scenario 1: Platform Dominance (Very High Probability)

Amazon wins the lawsuit. Google, Apple, and other major platforms watch carefully and implement similar policies. The outcome:

  • Platforms allow only "approved" agents
  • Approved agents must share 15-30% revenue with platforms
  • Platforms build superior first-party agents using proprietary data
  • Market concentration increases dramatically

This is the most likely outcome because platforms hold all the leverage. They control access to inventory, customer data, and the ability to transact. If you want your AI agent to work, you play by their rules or you don't play at all.

Winner: Existing platform giants. The "disruption" looks suspiciously like the old oligopoly, just with AI agents instead of apps.

Scenario 2: Payment Network Mediation (Medium probability)

Visa and Mastercard successfully establish themselves as neutral trust brokers. Their authentication standards become mandatory. Multiple agents can compete, but all must register with payment networks and follow their protocols.

This creates a more open ecosystem than Scenario 1, but you've still got gatekeepers—just different ones. Every transaction generates payment network fees. The rails change hands, but someone still controls the rails.

Winner: Payment networks become infrastructure monopolies. Better than platform domination, but not exactly a free market.

Scenario 3: Regulatory Intervention (Very low)

Governments step in, mandate open access standards, require algorithmic transparency, and force interoperability. The EU tries this first with AI Act enforcement.

Winner: Consumers and smaller players benefit from enforced competition.

Reality check: Given current U.S. regulatory momentum and the fact that legal frameworks are years behind AI development, this seems highly unlikely. The platforms are moving too fast, and regulators are too slow.

Why This Probably Makes Markets Less Competitive

Here's the uncomfortable truth: despite all the talk about AI "democratizing" commerce and creating more efficient markets, the likely outcome is increased market concentration.

Why? 

Trust Concentrates Around Scale

When AI agents are making autonomous purchases with your money, you need to trust them completely. That trust is hard to build and easy to destroy. Large, established players like Amazon can credibly say "we've processed billions of transactions, here's our security track record."

A startup building a shopping agent? Much harder sell. The trust moat actually gets deeper, not shallower.

Data Moats Become too big wall to jump

The best shopping agent needs to know:

  • Your purchase history
  • Your preferences and budget
  • Your calendar and schedule
  • Your payment methods and addresses
  • Context about why you're shopping

Amazon already has all of this. Google has most of it. A third-party agent has... whatever you manually tell it.

This isn't a gap you can close with "better AI." It's a fundamental data disadvantage that compounds over time.

Network Effects Intensify

Just as traditional commerce requires an ecosystem (platforms, payment processors, logistics, fraud prevention), agentic commerce needs an even more complex interconnected system. The platforms that can bundle these services—authentication, payments, fulfillment, customer service—win by default.

It's the AWS playbook: provide the full stack, make integration seamless, and competitors can't match the convenience.

Power to Block Is Power to Control

This is the key insight from the Amazon-Perplexity fight: if platforms can simply block agents they don't like, then innovation requires permission.

Want to build a revolutionary shopping agent? Great. But if Amazon, Google, and Walmart all block you, your revolutionary agent can't access any inventory. You've built a car with no roads to drive on.

The platforms learned from the app store wars: let a thousand flowers bloom, then harvest the ones that matter.

What This Means for Engineers Building AI Applications

If you're working on AI agents, here's what you need to understand:

Platform Risk Is Your Existential Risk

Don't build on platforms you don't control unless you have explicit agreements in place. The terms of service you're operating under were written before agentic AI existed, and platforms can change the rules whenever they want.

Perplexity is learning this the hard way. They built a business model that required access to Amazon's platform, then discovered Amazon could just say "no."

The Liability Problem Won't Solve Itself

Right now, there's massive ambiguity about who's responsible when AI agents screw up. This ambiguity is risk for everyone in the stack. You need to:

  • Get explicit terms in writing about agent behavior and limits
  • Build audit trails for every decision your agent makes
  • Have clear escalation paths when things go wrong
  • Understand you're probably liable for your agent's actions, even if that's not fair

Security Can't Be an Afterthought

The threat model for agentic commerce is genuinely novel. You can't just apply traditional bot detection because legitimate agents are bots. You need:

  • Cryptographic agent authentication (like Web Bot Auth)
  • Behavioral anomaly detection that works for non-human actors
  • Multi-party verification for high-value transactions
  • Fallback to human-in-the-loop when confidence is low

This is hard, expensive, and essential. The first major security breach involving agent-based shopping will tank consumer trust in the entire category.

Think in Systems, Not Just Models

The failure mode for agentic commerce isn't "the AI makes a mistake." It's "the AI makes a reasonable-seeming decision based on incomplete data, which cascades into a mess of returns, chargebacks, and customer service nightmares."

Good agentic systems need:

  • Clear boundaries on what decisions they can make autonomously
  • Confidence thresholds that trigger human review
  • Graceful degradation when uncertain
  • Mechanisms for users to understand and override decisions

This is systems engineering, not just prompt engineering.

So Who Wins?

If you're asking "who wins the agentic commerce war," here's my read:

Tier 1: Platform Oligarchs (Amazon, Google, Apple) - They control access to inventory and customers. They can block competitors and extract rent from those they allow. Amazon's lawsuit against Perplexity is them establishing this reality.

Tier 2: Payment Networks (Visa, Mastercard) - Becoming the critical trust infrastructure. Every transaction flows through them, and they're setting authentication standards for the entire ecosystem.

Tier 3: AI Insurgents (OpenAI, Perplexity, Anthropic) - High risk, high reward. They have the AI capabilities and consumer mindshare, but they need platform access to deliver value. Many will get squeezed or forced into revenue-sharing deals.

The Losers: Traditional retailers and brands who get reduced to "background utilities" in agent-controlled marketplaces. TripAdvisor already down 30% in traffic. AllRecipes lost 15%. This is the canary in the coal mine.

The uncomfortable parallel: this is the app store model all over again. Platforms create "open" ecosystems, welcome innovation, then monetize, control, and eventually squeeze everyone building on top.

In five years, we'll have agentic commerce. But it will likely be dominated by 3-5 massive platforms that control access, set standards, and extract rent. The "revolution" will look suspiciously like the old regime—just with better AI.

The Bottom Line

Agentic commerce is coming whether we're ready for it or not. The technology works, the market opportunity is massive, and the big platforms are already building it.

But let's not fool ourselves about what we're building. This isn't some perfect future where AI agents create perfect market efficiency and infinite consumer choice. It's a new battleground for the same old fight: who gets to control access to customers, and who gets to extract rent from transactions.

Amazon is suing Perplexity because they understand what's at stake. This isn't about "covert access" or "customer security"—those are the legal justifications. The real fight is about whether Amazon gets to control agentic commerce the same way Apple controlled app distribution and Google controlled digital advertising.

And based on history, platform power, and the economics of trust at scale, they probably will.

We've seen this movie before. The technology is new, but the plot is depressingly familiar.