GitHub Taskflow Agent Finds 24 Android Vulnerabilities With AI-Guided Audits
GitHub Security Lab reported on September 28 that its open-source Taskflow Agent and Android-specific audit workflows have identified 24 vulnerabilities in Android applications. The work extends GitHub's earlier AI-assisted security research into mobile-specific attack surfaces, using structured taskflows to guide language models through entry-point discovery, vulnerability classification and report generation.
GitHub has published the agent framework and taskflows, documented how researchers structure the audits, and provided a repeatable workflow that other security teams can apply to their own repositories. Model findings still require human security review because GitHub's testing found errors in severity estimates and exploitability judgments.
What GitHub changed for Android auditing
GitHub Security Lab's Taskflow Agent is an MCP-enabled framework for declarative, YAML-driven agent workflows. Its taskflows break an audit into smaller model tasks instead of asking one model to reason about an entire application in a single prompt.
For the Android work, GitHub added a taskflow that separates mobile entry points from other application entry points. It also expanded application classification prompts with mobile vulnerability classes such as confused-deputy issues, insecure broadcasts, WebView problems and exposed JavaScript bridges.
The researchers run complementary prompts multiple times. One constrains the model toward known mobile vulnerability patterns, while a broader pass gives it room to identify relationships that a narrowly specified checklist could miss. GitHub says this combination helped surface both straightforward flaws such as path traversal and higher-impact application issues.
24 findings, with human validation still required
GitHub says the Android taskflows had found 24 vulnerabilities at the time of publication. The research includes examples involving cross-application scripting in WebViews and exposed JavaScript bridges, alongside lower-severity issues.
The same research identifies severity assessment as a weak point for the model-driven workflow. A model can flag a technically suspicious path while missing a mitigating condition elsewhere in the application. GitHub gives the example of storage precedence changing whether an apparent path-traversal issue has real impact.
The system can search broadly, assemble candidate findings and produce evidence for a researcher. Human reviewers then validate reachability, application state, mitigations and severity before disclosure.
How the published workflow runs
GitHub's Android audit instructions use the open-source seclab-taskflows workflow from a Codespace. The documented mobile audit entry point is:
./scripts/audit/run_mobile.sh myorg/myrepo
GitHub says a medium-sized repository can take roughly one to two hours. Results are written to an SQLite database, where the audit_results table identifies candidate vulnerabilities for review.
A GitHub Copilot license is required for the documented workflow, and runs consume premium model requests. GitHub warns that the taskflows can generate many tool calls and substantial token usage. The underlying Taskflow Agent itself requires Python 3.10 or newer when deployed from source and is licensed under MIT.
Why structured taskflows matter
The architecture addresses a recurring problem in AI code review: large, loosely defined prompts can omit steps or lose track of application-specific context. Taskflows encode individual stages and dependencies in YAML, allowing researchers to refine one part of the audit without rewriting the entire process.
GitHub's framework can also connect agents to MCP tools and CodeQL-oriented workflows. Earlier Security Lab work used taskflows to triage CodeQL alerts and reported roughly 30 real-world vulnerabilities from that process. In March, GitHub said broader auditing taskflows had reported more than 80 vulnerabilities, with about 20 publicly disclosed at that time.
The Android results provide a focused test of the approach because mobile applications have platform-specific entry points, permission models and inter-component behavior. The published methodology shows that domain-specific prompts and repeated passes can materially change what the model examines.
Deployment considerations
Teams evaluating the workflow should budget for model usage, repeated runs and researcher review. GitHub explicitly recommends multiple runs because model behavior is non-deterministic, and its researchers manually validate findings before reporting them.
For security engineering teams, the framework provides an inspectable implementation of the agent, taskflow definitions and methodology. Teams can adapt the prompts to their own threat models and record project-specific checks that reduce false positives over time.
Sources
- GitHub Security Lab, How we found 24 Android vulnerabilities using our open source AI security agent (September 28, 2026): https://github.blog/security/how-we-found-24-android-vulnerabilities-using-our-open-source-ai-security-agent/
- GitHub Security Lab Taskflow Agent repository: https://github.com/GitHubSecurityLab/seclab-taskflow-agent
- GitHub Security Lab, How to scan for vulnerabilities with GitHub Security Lab's open source AI-powered framework (March 2026): https://github.blog/security/how-to-scan-for-vulnerabilities-with-github-security-labs-open-source-ai-powered-framework/
- Help Net Security independent coverage (September 29, 2026): https://www.helpnetsecurity.com/2026/09/29/github-ai-android-app-vulnerabilities/