Introduction

AI coding agents are increasingly being used for more than code generation. They can inspect repositories, reason about application behavior, run tools, analyze test results, and investigate security issues.

GitHub has described an experiment in which an AI-powered security agent was used to investigate Android applications and discovered 24 vulnerabilities. The work provides an interesting example of how AI agents can be applied to security research rather than only software development.

The important engineering lesson is not simply that an AI agent found vulnerabilities. The more useful question is how an agent can combine repository analysis, security reasoning, testing, and iteration into a repeatable vulnerability-discovery workflow.

This article explains that workflow, where AI security agents can help, where human security expertise remains necessary, and how development teams can apply similar ideas to their own applications.

What Is an AI Security Agent?

A traditional AI coding assistant generally responds to a developer's request.

An AI security agent can operate with a broader objective.

Instead of asking:

Explain this function.

A security-focused agent might receive a goal such as:

Investigate this Android application for security vulnerabilities.

The agent can then work through multiple steps:

Repository
    |
    v
Understand Application
    |
    v
Identify Attack Surfaces
    |
    v
Generate Hypotheses
    |
    v
Inspect Code
    |
    v
Run Tests / Tools
    |
    v
Validate Finding
    |
    v
Document Evidence

This makes the agent more similar to an automated security researcher than a traditional autocomplete system.

How AI Agents Change Security Testing

Conventional security testing often involves several separate activities:

  • Source-code review

  • Static analysis

  • Dependency analysis

  • Dynamic testing

  • Manual investigation

  • Vulnerability validation

  • Security reporting

An AI agent can coordinate several of these activities.

For example:

AI Agent
   |
   +---- Search source code
   |
   +---- Inspect configuration
   |
   +---- Trace data flow
   |
   +---- Identify suspicious behavior
   |
   +---- Run available tools
   |
   +---- Review results
   |
   +---- Investigate further
   |
   +---- Produce evidence

The value comes from connecting these steps rather than performing only one automated scan.

Why Android Applications Are Interesting Targets

Android applications expose several different attack surfaces.

An application may contain:

  • Activities

  • Services

  • Broadcast receivers

  • Content providers

  • Deep links

  • WebViews

  • Intents

  • Local storage

  • Network communication

  • Authentication logic

  • Native libraries

  • Exported components

A security investigation therefore requires understanding how different parts of the application interact.

For example, an exported Android component may appear harmless when examined independently.

The security risk may only become visible when the component:

  1. Accepts attacker-controlled input.

  2. Passes that input to another component.

  3. Uses the value in a sensitive operation.

  4. Returns information to an unauthorized caller.

An agent capable of tracing this chain can investigate beyond individual lines of code.

The Vulnerability Discovery Workflow

A useful AI security workflow starts with reconnaissance.

The agent first builds an understanding of the application.

Step 1: Understand the Repository

The agent can inspect:

AndroidManifest.xml
Gradle configuration
Java/Kotlin source
Resources
Network configuration
Native libraries
Tests

The goal is to understand the application structure before searching for vulnerabilities.

This avoids treating every suspicious-looking line as a security issue.

Step 2: Identify Attack Surfaces

Next, the agent can identify externally reachable components.

For Android applications, these may include:

Activities
Services
Receivers
Providers
Deep links
WebViews
Network endpoints

The agent can then prioritize areas where untrusted input crosses a security boundary.

Step 3: Generate Security Hypotheses

Instead of scanning blindly, the agent can form hypotheses.

For example:

Could an exported Activity expose sensitive functionality?

Can an external Intent control a file path?

Does a WebView accept untrusted content?

Can a deep link bypass authentication?

Does an application store credentials insecurely?

Each hypothesis can then be investigated.

This is closer to human security research than a simple pattern-matching scanner.

Step 4: Trace Data Flow

Suppose an application receives a value through an Intent.

The agent can investigate:

External Intent
      |
      v
Intent.getStringExtra()
      |
      v
Application Method
      |
      v
File Operation
      |
      v
Sensitive Resource

The important question is whether untrusted input can influence the sensitive operation.

Data-flow reasoning is particularly valuable because many vulnerabilities are caused by interactions between otherwise normal pieces of code.

Step 5: Validate the Finding

Finding suspicious code is not enough.

The agent should attempt to establish whether the behavior is actually exploitable.

A useful validation process is:

Suspicious Code
      |
      v
Security Hypothesis
      |
      v
Reproduction Attempt
      |
      v
Observed Behavior
      |
      v
Evidence

This reduces the chance of reporting every static-analysis warning as a confirmed vulnerability.

What Makes Agent-Based Security Different?

Traditional security automation usually follows predefined rules.

For example:

IF exported component
AND sensitive operation
THEN report finding

An agent can instead perform iterative investigation:

Find component
      |
Inspect code
      |
Understand input
      |
Trace execution
      |
Find protection
      |
Determine whether protection can be bypassed
      |
Test hypothesis

This allows the investigation to adapt based on what the agent discovers.

The 24-Vulnerability Result

GitHub reported that its security research using AI agents uncovered 24 vulnerabilities in Android applications.

The result demonstrates the potential of combining AI reasoning with security tooling and repository-level analysis.

However, the number itself should not be interpreted as evidence that AI agents can independently replace professional security researchers.

Vulnerability discovery depends heavily on:

  • Target selection

  • Agent capabilities

  • Tool access

  • Prompting and task design

  • Validation methodology

  • Human review

  • Application complexity

The meaningful lesson is the workflow used to investigate the applications.

AI Agent vs Traditional Static Analysis

Capability

Static Analysis

AI Security Agent

Pattern matching

Strong

Strong

Repository exploration

Limited by rules

Flexible

Hypothesis generation

Limited

Strong potential

Data-flow reasoning

Rule-dependent

Can reason across code

Tool orchestration

Usually predefined

Can coordinate multiple tools

Iterative investigation

Limited

Strong

False-positive investigation

Rule-dependent

Can investigate context

Human review

Still required

Still required

This does not mean one approach replaces the other.

Static analysis provides repeatable automated checks. AI agents can help investigate findings and explore paths that are harder to express as fixed rules.

Combining AI Agents With Security Tools

An effective security agent should not operate only on source code.

It can potentially coordinate tools such as:

Repository Search
      |
      +---- Static Analysis
      |
      +---- Dependency Scanner
      |
      +---- Build System
      |
      +---- Test Runner
      |
      +---- Emulator
      |
      +---- Network Analysis
      |
      +---- Security Database

The agent can use tool results as evidence for its next investigation step.

For example:

Static Analysis
      |
      v
Suspicious API Usage
      |
      v
Agent Searches Callers
      |
      v
Finds External Input
      |
      v
Runs Test
      |
      v
Confirms Behavior

This is one of the most interesting aspects of agent-based security research.

Why Validation Matters

AI models can produce convincing explanations even when their conclusions are incorrect.

Security workflows therefore need evidence.

A useful finding should answer:

  1. What is vulnerable?

  2. Where is the vulnerable code?

  3. What input controls the behavior?

  4. What security boundary is crossed?

  5. Can the behavior be reproduced?

  6. What impact does it have?

  7. How can the issue be fixed?

A report that only says:

This code may be vulnerable.

is not enough.

A stronger finding provides the reasoning and evidence necessary for a security engineer to reproduce and evaluate the issue.

Human Security Researchers Still Matter

AI agents can accelerate investigation, but human expertise remains important.

Security researchers provide:

  • Threat modeling

  • Business-context understanding

  • Risk assessment

  • Exploit validation

  • Responsible disclosure

  • Vulnerability classification

  • Remediation guidance

  • Adversarial thinking

An AI agent may discover a suspicious behavior, but determining its real-world impact can require knowledge outside the repository.

For example, an application may expose an API internally, but the security significance depends on authentication, deployment architecture, data sensitivity, and surrounding controls.

Using AI Security Agents in Development

Development teams can apply similar concepts without turning every developer workstation into a penetration-testing environment.

A controlled workflow might look like:

Pull Request
     |
     v
Automated Security Checks
     |
     v
AI Investigation
     |
     +---- No significant finding
     |
     +---- Potential finding
              |
              v
         Human Review
              |
              v
        Remediation

This allows AI to act as an additional investigation layer.

Security Agent Guardrails

Security agents require stronger controls than ordinary coding assistants.

A production implementation should consider:

Repository Access

Give the agent access only to repositories it needs.

Tool Permissions

Separate read-only analysis from actions that modify code or infrastructure.

Network Access

Control outbound network access carefully.

Secrets

Never expose production credentials unnecessarily.

Evidence

Require findings to include reproducible evidence.

Human Approval

Require human review before high-impact actions.

A useful permission model is:

Level 1
Read repository
    |
    v
Level 2
Run tests
    |
    v
Level 3
Run security tools
    |
    v
Level 4
Create proposed patch
    |
    v
Level 5
Human-approved change

The agent should not automatically receive production privileges simply because it can analyze application code.

Common Mistakes

Treating AI Findings as Confirmed Vulnerabilities

An AI-generated finding is a hypothesis until it has been validated.

Asking for a Generic Security Review

A broad request such as:

Find security problems.

may produce less useful results.

Specific objectives produce better investigations.

For example:

Inspect exported Android components and determine whether external callers can reach privileged functionality without authentication.

Ignoring Application Context

Source code alone does not always reveal the complete security boundary.

Giving Excessive Permissions

A security agent should have the minimum permissions required for its task.

Skipping Reproduction

A convincing explanation is not the same as a reproducible vulnerability.

Troubleshooting AI Security Investigations

The Agent Finds Too Many Issues

Narrow the investigation scope.

Start with one attack surface such as exported components, WebViews, authentication, or insecure storage.

Findings Are Repetitive

Ask the agent to group findings by root cause and identify duplicate manifestations.

Findings Lack Evidence

Require source locations, data-flow paths, reproduction steps, and observed behavior.

The Agent Gets Stuck

Break the investigation into smaller tasks.

For example:

Analyze manifest
      |
Analyze exported components
      |
Trace external inputs
      |
Investigate sensitive operations

This is often easier to evaluate than one large security request.

Best Practices

  1. Use AI agents as investigation assistants, not autonomous security authorities.

  2. Give agents narrowly defined security objectives.

  3. Start with repository reconnaissance.

  4. Prioritize externally reachable attack surfaces.

  5. Combine source analysis with runtime testing where appropriate.

  6. Require evidence for security findings.

  7. Validate important findings manually.

  8. Keep agent permissions minimal.

  9. Protect credentials and sensitive repository data.

  10. Separate analysis privileges from code-modification privileges.

  11. Record investigation steps for reproducibility.

  12. Integrate AI investigation with existing security tooling.

  13. Treat remediation suggestions as proposed changes that require review.

  14. Measure false positives as well as vulnerabilities discovered.

Advantages and Disadvantages

Advantages

  • Can explore large repositories quickly

  • Can connect information across multiple files

  • Can generate and test security hypotheses

  • Can coordinate multiple security tools

  • Can reduce repetitive investigation work

  • Can help developers understand security findings

  • Can accelerate vulnerability research

Disadvantages

  • AI-generated findings can be incorrect

  • Validation remains necessary

  • Complex security context can be difficult to infer

  • Tool permissions require careful design

  • Security investigations can consume significant compute and runtime resources

  • Agents may miss vulnerabilities outside their investigation scope

  • Sensitive code and data require strict access controls

Where AI Security Agents Fit

AI security agents are particularly useful for tasks that require investigation rather than simple pattern matching.

Examples include:

  • Exploring large repositories

  • Investigating static-analysis findings

  • Tracing untrusted input

  • Reviewing authentication paths

  • Analyzing exported Android components

  • Investigating WebView behavior

  • Examining dependency-related risks

  • Building reproduction cases

  • Preparing security reports

They work best as another layer in a broader security program.

Summary

GitHub's report of 24 Android vulnerabilities demonstrates how AI agents can be used for security research beyond traditional code generation.

The important capability is the agent's ability to investigate. It can inspect a repository, identify attack surfaces, generate hypotheses, trace data flow, use security tools, test suspicious behavior, and collect evidence.

That workflow can complement static analysis and conventional security testing, particularly when vulnerabilities depend on interactions across multiple components.

AI security agents should still operate under strong guardrails. Findings need validation, permissions should be limited, sensitive information must be protected, and human security expertise remains essential for assessing impact and remediation.

For development teams, the practical opportunity is to use AI agents as an additional security investigation layer, helping researchers and developers spend more time on meaningful vulnerabilities and less time on repetitive analysis.