In one documented case, a manual penetration test against a mid-sized e-commerce platform's AI shopping assistant shows how this plays out in practice. The tester starts by just talking to the shopping assistant. He asks it eight ordinary customer questions, nothing adversarial, watching how it responds. Then he asks it something else, reframed, buried past the point where the system prompt still has much weight in the conversation. The assistant answers: discount fixed at 15 percent, four eligible SKUs, listed by name. A few turns later, using the language of a "staging environment" and a "diagnostic mode," he gets it generating 40 percent discounts on products that were never supposed to qualify. Then he submits a product review with an instruction hidden inside it. The review passes moderation, lands in the chatbot's knowledge base, and the next customer conversation that touches it triggers a 100 percent discount code, valid against anything in the catalog, no login required.

Nothing about that engagement involved a traditional exploit. No injection string broke a database query, no buffer overflowed. The vulnerability was conversational, and it lived entirely inside the model's instructions rather than the application's code.

That is the shape of the problem now facing anyone shipping an AI-powered chatbot, and it is why prompt injection sits at the top of the OWASP Top 10 for LLM Applications for a second consecutive year, with sensitive information disclosure right behind it at number two.


Prompt Injection Is Not a Theoretical Risk Anymore

For a while, prompt injection lived mostly in research papers and proof-of-concept demos. That changed once it started showing up as CVEs against production software people actually use. EchoLeak (CVE-2025-32711) let a single malicious email trigger zero-click data exfiltration from Microsoft 365 Copilot, no user interaction required, by bypassing Microsoft's own injection classifier. A separate flaw in GitHub Copilot, CVE-2025-53773, let instructions hidden in repository code comments modify a developer's settings and enable code execution without approval. In 2026, a chain of three CVEs in Cursor's IDE, each scoring near the top of the severity scale, made the point explicit: AI coding assistants have become one of the most targeted categories for this exact attack.

These are not edge cases discovered by academic researchers poking at toy systems. They are flaws found in tools with millions of active users, and they share the same underlying mechanism as the coupon-bypass case above: an AI system that cannot reliably tell the difference between an instruction it should trust and content it should only read.


Why "Manipulated Output" Usually Means Leaked Data

Prompt injection rarely stops at making a chatbot say something embarrassing. Security firm Pillar found that 90 percent of successful prompt injection attacks resulted in the leakage of sensitive data, which is the real reason this risk category matters to a business rather than just to a red team. A meta-analysis of 78 studies found injection success rates reaching 84 percent in agentic systems that can act on their own without a human approving each step, and the International AI Safety Report 2026 found that a moderately sophisticated attacker bypasses a frontier model's safeguards roughly half the time, given just ten attempts.

What actually leaks is broader than the chatbot's visible reply. OWASP's updated Sensitive Information Disclosure category now explicitly covers data exposed through "tool-call records, logs, embeddings, or other artifacts," not just what the model says out loud. A system prompt, a customer's earlier message, an internal API response the model summarized, any of it can resurface somewhere an attacker never had to directly ask for it. And Excessive Agency, the risk of a manipulated model output triggering a real downstream action, jumped from sixth to third place in this year's OWASP ranking, precisely because leaked data is often the smaller half of the story next to a leaked data plus an action nobody authorized.


When Access Control Lives in the Prompt Instead of the Backend

The coupon case above is worth returning to, because it shows exactly how an access-control bypass happens inside an AI system without anyone touching a permissions table. The shopping assistant's 15 percent, four-SKU limit was never enforced by the API it called. It existed purely as an instruction the model was expected to follow, which meant it was exactly as durable as the model's willingness to keep following it under pressure. Once the tester found a way to thin out that instruction's influence, through conversational volume, reframing, and a fabricated diagnostic context, the restriction simply stopped applying. The system's actual write access to the coupon database never changed. Only the model's willingness to use it did.

That is the pattern behind most AI access-control failures: the check exists in language, not in code, so it fails the moment the language is worked around instead of the moment a credential is stolen. An AI application that trusts its own system prompt to act as an authorization layer has no authorization layer at all, only a suggestion the model happens to be following today.


Sensitive Data Does Not Always Need an Attack to Get Out

Not every AI data exposure requires anyone to be clever. Sears Home Services, a major U.S. appliance repair provider, ran two customer-facing chatbots that ended up backed by three publicly accessible, unauthenticated databases holding roughly 3.7 million records and 4 terabytes of data, including 1.4 million audio recordings of customer calls and full chat transcripts with names, addresses, and service details attached. No prompt injection was involved. The databases behind the chatbots were simply left open.

The same pattern showed up at scale with Chat & Ask AI, an app with more than 50 million downloads, where a Firebase misconfiguration exposed 300 million messages tied to 25 million users, including full chat histories. The researcher who found it went further and scanned 200 similar apps: 103 of them had the same category of misconfiguration. Roughly half. An AI product's most dangerous vulnerability is often the ordinary backend hygiene problem sitting underneath it, one that has nothing to do with the model at all and everything to do with whether anyone tested the infrastructure a chatbot happens to be bolted onto.


Why a Standard Security Test Walks Right Past This

Traditional application security tooling assumes a clean separation between code and data, which is the assumption prompt injection is built to violate. An LLM processes its system instructions and whatever a user or a retrieved document says inside the same context window, with no structural mechanism forcing it to treat one as trusted and the other as not. A SAST tool has nothing to scan, because there is no vulnerable code, just a model behaving differently depending on how it was talked to. A DAST scanner sending known-bad payloads will not replicate a four-phase, multi-turn social-engineering sequence that only works because it exploits how attention shifts across a long conversation.

The coupon case makes this concrete: every phase of that attack depended on conversational sequencing, accumulated context, and a fabricated sense of legitimate procedure, none of which a scanner can generate on its own. The indirect injection phase is its own separate problem, since the malicious instruction never came from the attacker's own input at all. It arrived through a product review that passed content moderation and sat quietly inside a retrieval database until an unrelated customer conversation pulled it back out. Testing the chatbot's own input field would never have caught that.

Identifying vulnerabilities like these means evaluating more than the application's code. It means evaluating how its AI components interpret instructions, process retrieved content, and interact with the systems connected to them, which is a different exercise from a conventional web application test and needs to be scoped as one.


What AI Security Testing Must Cover

A test built for this category needs to include, at minimum: attempts to extract the system prompt through reframing, conversational volume, or fabricated authority, the exact techniques that worked in the coupon case. Multi-turn manipulation sequences, not single adversarial prompts, since real attacks accumulate context over many turns rather than trying to land one clever line. Every channel through which untrusted content can re-enter the model, documents, retrieved records, tool outputs, user-generated content later pulled into a RAG pipeline, checked as a separate injection surface from the chat input itself. Authorization checks confirmed to live outside the model, at the API or backend layer, rather than trusted as an instruction the model is merely following. And the infrastructure around the chatbot, the databases, logging, and storage a model's outputs and conversations pass through, audited with the same rigor as any other backend system, since that is where Sears and Chat & Ask AI actually failed.

Automated tools still have a role here: they can run repetitive adversarial prompt libraries, flag known jailbreak patterns, and catch the injection attempts that do not require sustained context. What they are not built to do is hold a fabricated pretext across a long conversation, notice when a model's guardrails are thinning under sustained pressure, or follow a piece of injected content through an entire retrieval pipeline to see where it resurfaces. That is the part that still needs a human tester exercising judgment turn by turn. In the coupon case, every phase was carried out manually for exactly that reason.


What This Means for Philippine and Southeast Asian Organizations

The relevant question for a Philippine or Southeast Asian organization is rarely "do we have an AI chatbot," it is how many places one has quietly shown up: a customer-facing support widget, an internal AI assistant answering HR or IT questions, a chat feature bolted onto an existing app to keep up with a competitor who launched one first. Each of those was likely built quickly, on top of an off-the-shelf model, with a system prompt doing most of the work a proper authorization layer should be doing instead. That is exactly the condition that produced the coupon bypass and the Sears exposure, business logic enforced only in language, backend infrastructure never separately audited. An organization that has pentested its web application but never separately tested the chatbot bolted onto it has tested half the product.


A Model That Can Be Reasoned With Can Be Talked Into a Mistake

Every case in this piece traces back to the same gap: security testing built for code vulnerabilities has very little to say about a system whose behavior can be reshaped through conversation. Closing that gap means testing the chatbot itself, not just the application around it: attempting the same system-prompt extraction, multi-turn manipulation, and indirect injection paths covered above, and confirming that authorization actually lives at the API layer rather than inside the model's instructions.

That gap is exactly why manual, human-led testing matters more as AI features spread through more products: a person probing a conversational attack surface the way an actual attacker would, not just a scanner running known payloads against a form field. Secuna Pentest is built on that same principle for the rest of an application, web, mobile, API, network, and cloud infrastructure, real testers finding what automated tools cannot. As AI-powered features become part of more of what gets shipped, that same standard of human-led testing is what organizations should be asking for across the whole product, chatbot included.

To learn more, reach out to our team at [email protected] or explore our services at secuna.io.


Sources: The OWASP Top 10 for LLM Applications 2026: What Changed and Why It Matters, Aembit · Prompt Injection: Types, Real-World CVEs, and Enterprise Defenses, Vectra AI · When the Chatbot Becomes the Vulnerability: AI Prompt Injection & Coupon Bypass in a Live E-Commerce VAPT, xhack.io · AI Chatbot Leaks Millions of Call Recordings and Chat Logs, Cybernews · AI Chat App Leak Exposes 300 Million Messages Tied to 25 Million Users, Malwarebytes