Introduction
AI coding agents are increasingly being used for more than code generation. They can inspect repositories, reason about application behavior, run tools, analyze test results, and investigate security issues.
GitHub has described an experiment in which an AI-powered security agent was used to investigate Android applications and discovered 24 vulnerabilities. The work provides an interesting example of how AI agents can be applied to security research rather than only software development.
The important engineering lesson is not simply that an AI agent found vulnerabilities. The more useful question is how an agent can combine repository analysis, security reasoning, testing, and iteration into a repeatable vulnerability-discovery workflow.
This article explains that workflow, where AI security agents can help, where human security expertise remains necessary, and how development teams can apply similar ideas to their own applications.
What Is an AI Security Agent?
A traditional AI coding assistant generally responds to a developer's request.
An AI security agent can operate with a broader objective.
Instead of asking:
Explain this function.
A security-focused agent might receive a goal such as:
Investigate this Android application for security vulnerabilities.
The agent can then work through multiple steps:
Repository
|
v
Understand Application
|
v
Identify Attack Surfaces
|
v
Generate Hypotheses
|
v
Inspect Code
|
v
Run Tests / Tools
|
v
Validate Finding
|
v
Document Evidence
This makes the agent more similar to an automated security researcher than a traditional autocomplete system.
How AI Agents Change Security Testing
Conventional security testing often involves several separate activities:
Source-code review
Static analysis
Dependency analysis
Dynamic testing
Manual investigation
Vulnerability validation
Security reporting
An AI agent can coordinate several of these activities.
For example:
AI Agent
|
+---- Search source code
|
+---- Inspect configuration
|
+---- Trace data flow
|
+---- Identify suspicious behavior
|
+---- Run available tools
|
+---- Review results
|
+---- Investigate further
|
+---- Produce evidence
The value comes from connecting these steps rather than performing only one automated scan.
Why Android Applications Are Interesting Targets
Android applications expose several different attack surfaces.
An application may contain:
Activities
Services
Broadcast receivers
Content providers
Deep links
WebViews
Intents
Local storage
Network communication
Authentication logic
Native libraries
Exported components
A security investigation therefore requires understanding how different parts of the application interact.
For example, an exported Android component may appear harmless when examined independently.
The security risk may only become visible when the component:
Accepts attacker-controlled input.
Passes that input to another component.
Uses the value in a sensitive operation.
Returns information to an unauthorized caller.
An agent capable of tracing this chain can investigate beyond individual lines of code.
The Vulnerability Discovery Workflow
A useful AI security workflow starts with reconnaissance.
The agent first builds an understanding of the application.
Step 1: Understand the Repository
The agent can inspect:
AndroidManifest.xml
Gradle configuration
Java/Kotlin source
Resources
Network configuration
Native libraries
Tests
The goal is to understand the application structure before searching for vulnerabilities.
This avoids treating every suspicious-looking line as a security issue.
Step 2: Identify Attack Surfaces
Next, the agent can identify externally reachable components.
For Android applications, these may include:
Activities
Services
Receivers
Providers
Deep links
WebViews
Network endpoints
The agent can then prioritize areas where untrusted input crosses a security boundary.
Step 3: Generate Security Hypotheses
Instead of scanning blindly, the agent can form hypotheses.
For example:
Could an exported Activity expose sensitive functionality?
Can an external Intent control a file path?
Does a WebView accept untrusted content?
Can a deep link bypass authentication?
Does an application store credentials insecurely?
Each hypothesis can then be investigated.
This is closer to human security research than a simple pattern-matching scanner.
Step 4: Trace Data Flow
Suppose an application receives a value through an Intent.
The agent can investigate:
External Intent
|
v
Intent.getStringExtra()
|
v
Application Method
|
v
File Operation
|
v
Sensitive Resource
The important question is whether untrusted input can influence the sensitive operation.
Data-flow reasoning is particularly valuable because many vulnerabilities are caused by interactions between otherwise normal pieces of code.
Step 5: Validate the Finding
Finding suspicious code is not enough.
The agent should attempt to establish whether the behavior is actually exploitable.
A useful validation process is:
Suspicious Code
|
v
Security Hypothesis
|
v
Reproduction Attempt
|
v
Observed Behavior
|
v
Evidence
This reduces the chance of reporting every static-analysis warning as a confirmed vulnerability.
What Makes Agent-Based Security Different?
Traditional security automation usually follows predefined rules.
For example:
IF exported component
AND sensitive operation
THEN report finding
An agent can instead perform iterative investigation:
Find component
|
Inspect code
|
Understand input
|
Trace execution
|
Find protection
|
Determine whether protection can be bypassed
|
Test hypothesis
This allows the investigation to adapt based on what the agent discovers.
The 24-Vulnerability Result
GitHub reported that its security research using AI agents uncovered 24 vulnerabilities in Android applications.
The result demonstrates the potential of combining AI reasoning with security tooling and repository-level analysis.
However, the number itself should not be interpreted as evidence that AI agents can independently replace professional security researchers.
Vulnerability discovery depends heavily on:
Target selection
Agent capabilities
Tool access
Prompting and task design
Validation methodology
Human review
Application complexity
The meaningful lesson is the workflow used to investigate the applications.
AI Agent vs Traditional Static Analysis
Capability | Static Analysis | AI Security Agent |
|---|---|---|
Pattern matching | Strong | Strong |
Repository exploration | Limited by rules | Flexible |
Hypothesis generation | Limited | Strong potential |
Data-flow reasoning | Rule-dependent | Can reason across code |
Tool orchestration | Usually predefined | Can coordinate multiple tools |
Iterative investigation | Limited | Strong |
False-positive investigation | Rule-dependent | Can investigate context |
Human review | Still required | Still required |
This does not mean one approach replaces the other.
Static analysis provides repeatable automated checks. AI agents can help investigate findings and explore paths that are harder to express as fixed rules.
Combining AI Agents With Security Tools
An effective security agent should not operate only on source code.
It can potentially coordinate tools such as:
Repository Search
|
+---- Static Analysis
|
+---- Dependency Scanner
|
+---- Build System
|
+---- Test Runner
|
+---- Emulator
|
+---- Network Analysis
|
+---- Security Database
The agent can use tool results as evidence for its next investigation step.
For example:
Static Analysis
|
v
Suspicious API Usage
|
v
Agent Searches Callers
|
v
Finds External Input
|
v
Runs Test
|
v
Confirms Behavior
This is one of the most interesting aspects of agent-based security research.
Why Validation Matters
AI models can produce convincing explanations even when their conclusions are incorrect.
Security workflows therefore need evidence.
A useful finding should answer:
What is vulnerable?
Where is the vulnerable code?
What input controls the behavior?
What security boundary is crossed?
Can the behavior be reproduced?
What impact does it have?
How can the issue be fixed?
A report that only says:
This code may be vulnerable.
is not enough.
A stronger finding provides the reasoning and evidence necessary for a security engineer to reproduce and evaluate the issue.
Human Security Researchers Still Matter
AI agents can accelerate investigation, but human expertise remains important.
Security researchers provide:
Threat modeling
Business-context understanding
Risk assessment
Exploit validation
Responsible disclosure
Vulnerability classification
Remediation guidance
Adversarial thinking
An AI agent may discover a suspicious behavior, but determining its real-world impact can require knowledge outside the repository.
For example, an application may expose an API internally, but the security significance depends on authentication, deployment architecture, data sensitivity, and surrounding controls.
Using AI Security Agents in Development
Development teams can apply similar concepts without turning every developer workstation into a penetration-testing environment.
A controlled workflow might look like:
Pull Request
|
v
Automated Security Checks
|
v
AI Investigation
|
+---- No significant finding
|
+---- Potential finding
|
v
Human Review
|
v
Remediation
This allows AI to act as an additional investigation layer.
Security Agent Guardrails
Security agents require stronger controls than ordinary coding assistants.
A production implementation should consider:
Repository Access
Give the agent access only to repositories it needs.
Tool Permissions
Separate read-only analysis from actions that modify code or infrastructure.
Network Access
Control outbound network access carefully.
Secrets
Never expose production credentials unnecessarily.
Evidence
Require findings to include reproducible evidence.
Human Approval
Require human review before high-impact actions.
A useful permission model is:
Level 1
Read repository
|
v
Level 2
Run tests
|
v
Level 3
Run security tools
|
v
Level 4
Create proposed patch
|
v
Level 5
Human-approved change
The agent should not automatically receive production privileges simply because it can analyze application code.
Common Mistakes
Treating AI Findings as Confirmed Vulnerabilities
An AI-generated finding is a hypothesis until it has been validated.
Asking for a Generic Security Review
A broad request such as:
Find security problems.
may produce less useful results.
Specific objectives produce better investigations.
For example:
Inspect exported Android components and determine whether external callers can reach privileged functionality without authentication.
Ignoring Application Context
Source code alone does not always reveal the complete security boundary.
Giving Excessive Permissions
A security agent should have the minimum permissions required for its task.
Skipping Reproduction
A convincing explanation is not the same as a reproducible vulnerability.
Troubleshooting AI Security Investigations
The Agent Finds Too Many Issues
Narrow the investigation scope.
Start with one attack surface such as exported components, WebViews, authentication, or insecure storage.
Findings Are Repetitive
Ask the agent to group findings by root cause and identify duplicate manifestations.
Findings Lack Evidence
Require source locations, data-flow paths, reproduction steps, and observed behavior.
The Agent Gets Stuck
Break the investigation into smaller tasks.
For example:
Analyze manifest
|
Analyze exported components
|
Trace external inputs
|
Investigate sensitive operations
This is often easier to evaluate than one large security request.
Best Practices
Use AI agents as investigation assistants, not autonomous security authorities.
Give agents narrowly defined security objectives.
Start with repository reconnaissance.
Prioritize externally reachable attack surfaces.
Combine source analysis with runtime testing where appropriate.
Require evidence for security findings.
Validate important findings manually.
Keep agent permissions minimal.
Protect credentials and sensitive repository data.
Separate analysis privileges from code-modification privileges.
Record investigation steps for reproducibility.
Integrate AI investigation with existing security tooling.
Treat remediation suggestions as proposed changes that require review.
Measure false positives as well as vulnerabilities discovered.
Advantages and Disadvantages
Advantages
Can explore large repositories quickly
Can connect information across multiple files
Can generate and test security hypotheses
Can coordinate multiple security tools
Can reduce repetitive investigation work
Can help developers understand security findings
Can accelerate vulnerability research
Disadvantages
AI-generated findings can be incorrect
Validation remains necessary
Complex security context can be difficult to infer
Tool permissions require careful design
Security investigations can consume significant compute and runtime resources
Agents may miss vulnerabilities outside their investigation scope
Sensitive code and data require strict access controls
Where AI Security Agents Fit
AI security agents are particularly useful for tasks that require investigation rather than simple pattern matching.
Examples include:
Exploring large repositories
Investigating static-analysis findings
Tracing untrusted input
Reviewing authentication paths
Analyzing exported Android components
Investigating WebView behavior
Examining dependency-related risks
Building reproduction cases
Preparing security reports
They work best as another layer in a broader security program.
Summary
GitHub's report of 24 Android vulnerabilities demonstrates how AI agents can be used for security research beyond traditional code generation.
The important capability is the agent's ability to investigate. It can inspect a repository, identify attack surfaces, generate hypotheses, trace data flow, use security tools, test suspicious behavior, and collect evidence.
That workflow can complement static analysis and conventional security testing, particularly when vulnerabilities depend on interactions across multiple components.
AI security agents should still operate under strong guardrails. Findings need validation, permissions should be limited, sensitive information must be protected, and human security expertise remains essential for assessing impact and remediation.
For development teams, the practical opportunity is to use AI agents as an additional security investigation layer, helping researchers and developers spend more time on meaningful vulnerabilities and less time on repetitive analysis.
Join the conversation! Your thoughts help the community grow.