Google says an internal artificial intelligence agent has identified more than 500 cross-site scripting vulnerabilities across the company’s first-party web applications.
Named PageBreak, the system combines AI-generated security hypotheses with separate validators that attempt to execute payloads against running applications. Google presents that validation stage as the mechanism that prevents plausible AI output from being mistaken for an exploitable vulnerability.
The reported total is substantial, but its scope requires care. It represents findings rather than affected users, and the available account does not establish that every reported flaw has been remediated. It also provides no CVE identifiers, CVSS scores, evidence of malicious exploitation or confirmed impact on Google customers.
PageBreak moved from a pilot to an internal product
PageBreak was developed by Google’s Product Security team to assess the company’s own web applications and browser extensions. Information security engineer Michał Bentkowski said it found more than 500 XSS flaws.
According to the report describing PageBreak and its results, the project began as a pilot last November and became a full product in January. The source does not give years for those two milestones. Google discussed the system in a company blog post last month.
These references describe development and disclosure milestones, not the dates on which individual vulnerabilities were introduced, discovered or fixed. There is also no single incident date because PageBreak is a security-testing system, not a reported breach.
Google’s public examples include three distinct findings:
- A cache-poisoning vulnerability affecting
apis.google.com - An XSS vulnerability in
admin.google.com - Insecure external handshakes involving browser extensions
Google’s companion post reportedly said those three issues had been fixed. That statement does not establish the remediation status of the wider collection, including the more than 500 XSS findings attributed to PageBreak.
The reporting also does not identify external organizations affected by the flaws. The stated testing scope covers Google’s first-party properties and browser extensions.
AI proposes vulnerabilities, but separate code must prove them
PageBreak’s architecture is intended to address a recurring limitation of AI-assisted security testing: a model can produce a technically convincing explanation without identifying a vulnerability that can actually be exploited.
Under the process described by Bentkowski, the agent first examines an application and generates a vulnerability hypothesis. It then passes that candidate to a specialized validator. The validator attempts to run a real payload against the application in an active environment, and PageBreak reports the issue only after that attempt demonstrates the vulnerability.
The validators are not written by AI, according to the description. Their logic and interfaces vary according to both the vulnerability category and the application surface being tested.
For example, XSS, SQL injection and remote code execution require different validation methods. Testing an HTTP application also presents a different interface from assessing a gRPC service. PageBreak therefore does not rely on one universal prompt or validation routine for every target.
This separation matters because the AI component does not have final authority over its own findings. It generates candidates, while another mechanism is responsible for producing execution evidence.
Google describes the system as combining autonomous discovery with deterministic validation. Bentkowski characterized its false-positive rate as “near-zero.” That figure remains a company-side claim: the available reporting does not include independent testing of PageBreak’s detection accuracy, false-positive rate or coverage.
Gemini models drive discovery, not final verification
PageBreak primarily uses Google’s Gemini 3.1 Pro and Gemini 3.5 Flash, although Google says the system can operate with other models.
The distinction between models and validators is central to the design. Gemini is used in the discovery process, where access to source code and security tooling can help the agent formulate potential attack paths. The non-AI validator then determines whether a candidate can be demonstrated against the running target.
Bentkowski said AI can discover vulnerabilities more quickly than human researchers. He also acknowledged the operational cost of sorting real weaknesses from credible-looking hallucinations. Without effective validation, faster discovery can simply transfer work to product security and engineering teams.
The available information does not provide comparative benchmarks against human researchers or other automated scanners. It also does not independently demonstrate the performance of either named Gemini model in this setting.
External experts support the separation of discovery and proof
Rickard Carlsson, CEO of Detectify, described a broader capacity problem facing security teams: AI tools can generate more suspected vulnerabilities than people can realistically investigate.
Carlsson argued that an agent should not verify its own conclusions. In his assessment, a separate validator should execute an actual payload against the running application before a proposed finding is accepted. That view aligns with PageBreak’s reported division between AI-led discovery and non-AI validation.
Darin Fredde, senior director of technical marketing engineering at Ridge Security, placed the approach within a shift toward continuous, autonomous and evidence-based offensive testing. He argued that discovery alone is insufficient. The process should continue through proof, remediation and verification that the repair works.
These comments are industry assessments rather than independent confirmation of PageBreak’s performance. Neither establishes how comprehensively the agent examined Google’s applications, how often it missed vulnerabilities or whether all validated findings were subsequently corrected.
Automated remediation remains a future step
Google intends to connect PageBreak with CodeMender, its automated vulnerability-remediation system. Under the proposed workflow, PageBreak would validate a flaw and CodeMender would generate a candidate code change. Product engineers would then review, validate and apply the proposed fix.
The integration is described as a future plan, not a currently operational remediation pipeline. Consequently, the more than 500 findings should not be interpreted as more than 500 automatically repaired vulnerabilities.
For Google’s product teams, PageBreak could reduce time spent manually reproducing AI-generated reports. A successful payload gives engineers stronger evidence that a suspected weakness deserves attention. It does not, by itself, establish severity, user impact or the correctness of a later patch.
For users and external administrators, the report supplies no product-specific patch instructions, indicators of compromise or workarounds. It also presents no evidence that attackers exploited the identified weaknesses. The immediate remediation responsibility therefore rests with Google and the teams maintaining the tested first-party services and extensions.
What PageBreak demonstrates, based on Google’s account, is a specific internal testing model: let an AI agent search broadly, but require separate executable proof before a result reaches engineers. Whether that architecture delivers its claimed low false-positive rate at scale has not been independently established.




