Original research · Statistical CTI · Published 29 August 2026

AI in Cyberattacks: A Statistical CTI Study of 111 Publications

Published · Last updated

A visualization-rich analysis of 103 eligible publications shows where AI appears in attacker workflows, which evidence is strongest, and which conclusions remain unsafe to make.

111 deduplicated publications103-publication denominator31 visualizationsEvidence-bounded
AI in Cyberattacks: A Statistical CTI Study cover showing a red and blue data stream forming a digital brain above research documents
116retrieved source records
111deduplicated publications
108usable indexed references
103primary analytical denominator

Artificial intelligence is now present in almost every discussion about cybercrime, state activity, fraud, phishing, malware, vulnerability research, and autonomous intrusion. The difficult analytical problem is no longer finding claims. It is separating observed attacker behavior from forecasts, controlled demonstrations, provider-abuse telemetry, vendor narrative, and repeated coverage of the same campaign.

This study converts 116 retrieved source records into 111 deduplicated publications and a normalized research dataset. The main statistical denominator is 103 publications assessed as relevant to attacker use of AI but still requiring analyst validation. Five contextual publications are kept outside the denominator, two broken source pages are excluded, and one non-AI report is excluded. The result is not an incident census. It is a map of what the collected evidence discusses, how often topics co-occur, and where the source base is mature or weak.

The companion interactive dashboard contains more than 30 widgets, heatmaps, ranked distributions, and a filterable publication explorer. The searchable References library links all 108 usable publications through normalized discovery facets. The underlying dataset README, CSV tables, SQLite database, and Excel workbook make every chart reproducible from the published normalized snapshot.

Scope and evidence boundary: The unit of analysis is a publication, not an incident, victim, intrusion, account, prompt, or malware sample. A publication can cover multiple incidents, and one incident can appear in multiple publications. All extracted tags, metrics, IOCs, actors, ATT&CK candidates, countries, sectors, CVEs, providers, and models require human validation against the canonical publisher source or the private working evidence record. At this published snapshot, all 111 analyst-review rows remain uncompleted. No chart proves attribution, successful exploitation, provider-specific abuse, or global prevalence.

Table of contents

  1. Executive findings
  2. Research question and evidence model
  3. Corpus construction and quality
  4. Temporal and source landscape
  5. How AI is described in cyberattacks
  6. Intrusion lifecycle and ATT&CK coverage
  7. Threat actors, targets, sectors, and geography
  8. LLM providers, models, and malicious-AI brands
  9. Attack vectors, infrastructure, data, and impact
  10. Quantitative metrics, IOCs, and CVEs
  11. Cross-dimensional analysis
  12. What defenders should do with these findings
  13. Explore the connected research ecosystem
  14. Limitations and reproducibility
  15. Conclusion
  16. References
  17. Follow My Work

Executive findings

The corpus produces seven defensible high-level findings:

  1. Reporting is concentrated in 2024–2026. Of 100 eligible publications with a known year, 46 are dated 2025. This is publication volume, influenced by collection scope and publisher activity; it is not a measured growth rate for attacks.
  2. AI is described across the full intrusion lifecycle. Actions on Objectives appears in 90 publications, Exploitation in 77, Delivery in 75, and Reconnaissance in 47. The breadth supports treating AI as a cross-workflow accelerator rather than a single technique.
  3. Identity and research dominate the use-case vocabulary. Identity fraud and impersonation appears in 45 eligible publications (43.7%), and reconnaissance/target research appears in 44 (42.7%). Obfuscation/evasion appears in 42 and malware development in 40.
  4. Initial Access is the leading ATT&CK tactic mention. It appears in 81 publications (78.6%), followed by Impact in 59 and Execution in 55. These are candidate mappings extracted from source text, not validated ATT&CK assignments for 81 distinct attacks.
  5. Government and financial services dominate sector mentions. Government appears in 53 publications (51.5%) and Financial Services in 48. This measures reporting attention and contextual references, not unique victims.
  6. Provider mentions are highly visible but easy to misread. OpenAI is named in 50 eligible publications, Anthropic in 33, and Google in 20. These counts include research, defensive discussion, provider telemetry, experiments, and product references. They do not mean 50 confirmed attacks used OpenAI.
  7. The corpus is rich in extractable observables but uneven in comparable outcomes. It contains 483 IOC candidate mentions and 919 unvalidated metric candidates. Only 29 eligible publications contain IOC candidates and 69 contain at least one quantitative candidate; the values use heterogeneous definitions and include extraction false positives, so they cannot be pooled as one loss, dwell-time, or blast-radius estimate.

Research dashboard

MeasureResultInterpretation
Archived source records116Downloaded HTML, PDF, or provenance-preserving proxy records
Deduplicated publications111Five companion-format duplicate groups collapsed
Main statistical denominator103Eligible with manual validation required
Context-only publications5Retained for background, excluded from main percentages
Excluded publications3Two broken pages and one non-AI source
Eligible tag mentions4,862Source-linked, machine-extracted occurrences; copied excerpts remain private
Tag dimensions23Actor, use case, sector, TTP, provider, impact, and other classes
Unique normalized tag values530Deduplicated within tag type
Unvalidated metric candidates919Percentages, duration, cost, blast-radius, and breakout-time strings; false positives remain
IOC candidate mentions483Hashes, defanged domains, and IPv4 strings requiring validation
Corpus disposition
Corpus disposition · publication coverage, not incident prevalence

Research question and evidence model

The central question is:

How does the collected 2022–2026 research describe AI use in attacker activity, and where does that evidence concentrate across use cases, intrusion stages, actors, targets, technologies, and measurable outcomes?

This is a descriptive corpus study. It does not estimate the fraction of all cyberattacks that use AI. There is no known sampling frame containing every attack or every report, and publication practices differ across governments, vendors, providers, researchers, and incident-response firms.

Unit of analysis

UnitUsed here?Meaning
PublicationYesOne deduplicated report, article, advisory, paper, or provider-abuse report
Source recordProvenance onlyOne downloaded representation; HTML/PDF companions can map to one publication
Incident or campaignNoNot consistently normalized across all sources
Tag co-mentionYesTwo normalized concepts appearing in the same publication
Victim or affected organizationNoNot consistently disclosed or uniquely countable
IOC candidateYes, separatelyA machine-extracted observable requiring analyst validation

Interpretation vocabulary

Corpus construction and quality

The private archive began with 116 retrieved source records. A content and metadata audit collapsed five companion-format pairs—HTML and PDF representations of the same publication—into 111 publication entities. The private source records retain the local files and extraction excerpts; the public provenance layer retains source IDs, canonical URLs, retrieval methods, hashes, duplicate decisions, content quality, and publication relationships without exposing archive paths.

The analytical dataset then records one row per publication and long-form tables for tags, metrics, and IOCs. The private working records retain evidence spans and excerpts so an analyst can move from a chart value back to the exact source context. The public exports retain source IDs, normalized values, confidence, extraction methods, and offsets, but intentionally omit copied excerpts and local archive paths.

Inclusion rules

Extraction quality is not claim validity

All 4,862 eligible tag occurrences are machine-extracted. High textual match confidence can show that a phrase exists near evidence, but it cannot prove the phrase describes a real malicious action, the named actor performed it, the provider enabled it, or the CVE was exploited with AI. Manual review remains mandatory.

Source types
Source types · publication coverage, not incident prevalence
Publisher concentration
Publisher concentration · publication coverage, not incident prevalence

The publisher distribution also reveals sampling bias. Google Threat Intelligence Group/Mandiant, Unit 42, Recorded Future, Check Point, CrowdStrike, OpenAI, and Anthropic contribute multiple records. Their visibility, disclosure policies, research priorities, and product telemetry shape the corpus.

Temporal and source landscape

Publication timeline
Publication timeline · publication coverage, not incident prevalence
YearPublicationsShare of eligible corpus
202211.0%
202387.8%
20242221.4%
20254644.7%
20262322.3%

The curve rises sharply through 2025, but it must not be read as an attack-growth curve. At least four alternative mechanisms can produce the same shape: greater publisher attention, more provider transparency reports, an expanding collection strategy, and real change in attacker experimentation. The corpus does not identify their separate effects. The 2026 count is also partial through the collection date, and three eligible publications have no resolved year.

How AI is described in cyberattacks

AI use cases

AI use cases
AI use cases · publication coverage, not incident prevalence
AI use casePublicationsShare of eligible corpus
Identity fraud and impersonation4543.7%
Reconnaissance and target research4442.7%
Obfuscation and evasion4240.8%
Malware development4038.8%
Deepfake video or image3635.0%
Translation and localization2827.2%
Vulnerability research2524.3%
Influence operations2120.4%
CAPTCHA bypass1716.5%
Deepfake voice1413.6%
Exploit development1413.6%
Phishing and lure generation1413.6%
Autonomous or agentic intrusion76.8%
Code debugging and scripting54.9%
Command and control21.9%

The leading pattern is not a single autonomous attack. It is the use of AI to reduce friction in familiar activities: identity impersonation, target research, evasion, malware development, localization, and vulnerability work. Autonomous or agentic intrusion appears in only seven eligible publications (6.8%), while identity fraud appears in 45 (43.7%) and target research in 44 (42.7%). This supports a bounded conclusion: the collected literature more often describes augmentation of existing attacker workflows than fully autonomous end-to-end compromise.

Use the References library to inspect the 45-publication identity-fraud slice rather than treating the aggregate as an incident count.

AI technology families

AI technologies
AI technologies · publication coverage, not incident prevalence

Generative AI and large language models dominate the technology vocabulary, while agentic AI appears in 38 publications and deepfake/synthetic media in 35. These categories overlap. A report can discuss an LLM-driven agent, generative malware assistance, and synthetic identity material simultaneously.

Intrusion lifecycle and ATT&CK coverage

Cyber Kill Chain

Kill Chain
Kill Chain · publication coverage, not incident prevalence

All seven Cyber Kill Chain phases are represented. Actions on Objectives, Exploitation, and Delivery are the most frequently covered. Weaponization is the least represented at 32 publications, but even that is present in nearly one-third of the eligible corpus. This breadth argues against designing “AI attack detection” as one alert category. Defenders need evidence across identity, email, endpoint, network, cloud, SaaS, and fraud telemetry.

ATT&CK tactic candidates

ATT&CK tactics
ATT&CK tactics · publication coverage, not incident prevalence

ATT&CK tactic tags are behavioral leads. They should be promoted into a defensive mapping only after the analyst validates the described behavior, actor, target, and evidence. Initial Access dominates because the corpus contains substantial social engineering, identity fraud, credential theft, vishing, phishing, and fake-site research.

The eligible source set is available as an Initial Access reference pivot, and the ATT&CK knowledge matrix provides the behavior-first context needed before operational mapping.

TTP families

TTP families
TTP families · publication coverage, not incident prevalence
TTP familyPublicationsShare of eligible corpus
Data exfiltration4846.6%
Command and control3735.9%
Deepfake impersonation3433.0%
Malware generation3332.0%
Defense evasion3231.1%
Credential theft2928.2%
Obfuscation2524.3%
Spearphishing2524.3%
Voice phishing / vishing2524.3%
PowerShell execution2322.3%
Supply-chain compromise2120.4%
Business email compromise1817.5%
Credential harvesting1817.5%
CAPTCHA bypass1716.5%
Prompt injection1615.5%

Data exfiltration is the leading normalized TTP family, appearing in 48 publications (46.6%). Command and control, deepfake impersonation, defense evasion, malware generation, and credential theft follow. These values measure corpus coverage, not the conditional probability that an AI-assisted intrusion will use each behavior.

Threat actors, targets, sectors, and geography

Named threat groups

Threat groups
Threat groups · publication coverage, not incident prevalence

APT28/Fancy Bear appears in 10 eligible publications, APT43/Kimsuky in nine, and Akira in seven. The names can describe attribution, comparison, historical background, or a source publisher’s own assessment. They must not be transformed into attribution facts without evidence-level review and alias resolution.

Sectors and target personas

Sectors
Sectors · publication coverage, not incident prevalence
Targets
Targets · publication coverage, not incident prevalence

Government, financial services, telecommunications, cryptocurrency, education, and critical infrastructure lead the sector vocabulary. Executives, developers, and employees lead the target-persona vocabulary. This combination is consistent with a research landscape focused on access, identity, code, financial abuse, and high-value social engineering. It does not establish unique victim counts or sector-specific attack rates.

Geography

Countries and regions
Countries and regions · publication coverage, not incident prevalence

Russia, China, the United States, North Korea, and Iran are the most frequently named countries in the eligible corpus. A country mention can refer to an actor’s alleged origin, a victim, infrastructure, law enforcement, a publication’s regional scope, or policy context. The country field is therefore unsuitable for a “most attacked country” ranking without relation-level coding.

LLM providers, models, and malicious-AI brands

Provider and model mentions

LLM providers
LLM providers · publication coverage, not incident prevalence
LLM models
LLM models · publication coverage, not incident prevalence

OpenAI appears in 50 eligible publications (48.5%), Anthropic in 33 (32.0%), and Google in 20 (19.4%). ChatGPT, Claude, and Gemini lead the product/model list. Provider-authored transparency reports are part of the corpus, so a provider can be counted because it published defensive findings, because another report discussed its service, or because the service appeared in an experiment. The correct label is provider mention coverage, not attacker market share. The OpenAI reference pivot exposes the eligible records behind that count.

Named malicious-AI tools

Malicious AI tools
Malicious AI tools · publication coverage, not incident prevalence

FraudGPT and WormGPT each appear in 12 publications. Underground names can be rebrands, scams, wrappers, marketing claims, or repeated secondary reporting. Their presence should trigger source and capability validation, not automatic classification as a distinct operational model.

Malware and tooling co-mentions

Malware and tooling co-mentions
Malware and tooling co-mentions · publication coverage, not incident prevalence

Named malware and tools are co-mentions in publications, not proof that AI generated, operated, or materially changed them. Use this distribution as a reading queue, then validate the exact claim and technical evidence through the Malware Analysis knowledge path.

Attack vectors, infrastructure, data, and impact

Attack vectors

Attack vectors
Attack vectors · publication coverage, not incident prevalence

Credential theft, vishing, spearphishing, supply-chain references, and business email compromise dominate. This supports a practical defensive priority: AI-related attack research must remain connected to established identity, email, browser, endpoint, and fraud controls rather than being isolated in a separate “AI” queue.

Infrastructure and data types

Infrastructure
Infrastructure · publication coverage, not incident prevalence
Data types
Data types · publication coverage, not incident prevalence

Email appears in 67 eligible publications, browsers in 61, and endpoints in 58. Credentials/passwords appear in 74 and documents/files in 65. These are publication co-mentions, but they identify high-value telemetry intersections for defensive validation.

Motivation and impact

Actor motivations
Actor motivations · publication coverage, not incident prevalence
Impacts
Impacts · publication coverage, not incident prevalence

Ransomware/extortion is the most frequent impact theme at 69 publications, followed by data theft/exfiltration at 53. Influence operations lead the motivation tags at 28 publications, followed by financial motivation at 21. Motivation is often a publisher inference; impact can be forecast, simulated, or observed. Analysts must retain that distinction.

Evidence maturity

Evidence landscape
Evidence landscape · publication coverage, not incident prevalence

Forecast or prediction appears in 52 eligible publications, incident response in 36, underground-market observation in 27, controlled study in 26, proof of concept in 20, and in-the-wild observation in 19. These categories overlap, but the mix shows why a single headline percentage would be misleading. The dataset combines prospective assessments, operational observations, controlled research, and provider telemetry.

For the most operational subset, start with the in-the-wild evidence pivot, then verify each publisher's evidence and methodology.

Quantitative metrics, IOCs, and CVEs

Quantitative metric coverage

Metric coverage
Metric coverage · publication coverage, not incident prevalence

The eligible corpus contains 919 machine-extracted metric candidates across 69 publications. Percentages are the dominant extracted class, but they describe different denominators: phishing success, fraud growth, observed actor activity, affected organizations, model performance, or other source-specific measures. The extraction also retains false-positive strings such as years, CVE fragments, and publisher boilerplate. Duration and blast-radius candidates are similarly heterogeneous. This study reports candidate coverage for review, not pooled means or confirmed outcomes.

IOC candidates

IOC composition
IOC composition · publication coverage, not incident prevalence

The IOC table contains 483 candidate mentions across 29 eligible publications. The largest categories are defanged domains and cryptographic hashes. These candidates must be validated for exact syntax, source context, temporal relevance, infrastructure ownership, and false-positive risk before blocking, enrichment, or sharing.

CVE mentions

CVEs
CVEs · publication coverage, not incident prevalence

The dataset contains explicit CVE strings only. A CVE mentioned in the same publication as AI does not prove AI-assisted vulnerability discovery or exploitation. Relation-level evidence would be required to make that claim.

Cross-dimensional analysis

Cross-dimensional heatmaps count publications containing both normalized tags. They are useful for triage and hypothesis generation, but they do not establish a direct semantic relation between the two tags.

Sector × AI use case

Sector by AI use case
Sector by AI use case · publication coverage, not incident prevalence

The heatmap highlights where reporting attention intersects. Identity, reconnaissance, evasion, malware development, and deepfake content recur across government, finance, telecommunications, cryptocurrency, and critical-infrastructure discussions.

Threat group × AI use case

Group by AI use case
Group by AI use case · publication coverage, not incident prevalence

This view is an analyst work queue. A high cell means several publications name the group and use case together. Each underlying claim still requires attribution, time, actor-alias, and behavior review against the canonical publisher source or private working evidence record.

Kill Chain × AI use case

Kill Chain by AI use case
Kill Chain by AI use case · publication coverage, not incident prevalence

The cross-tab shows that the most common AI functions span multiple phases. Reconnaissance and identity content can support access preparation, while malware development and evasion can appear from weaponization through installation and objectives.

Provider × AI use case

Provider by AI use case
Provider by AI use case · publication coverage, not incident prevalence

Provider/use-case co-mentions are especially sensitive to overinterpretation. Provider reports often describe abuse across many categories, which increases both row and column counts. Use the cells to locate evidence, never as a provider abuse-rate comparison.

Sector × geography

Sector by country
Sector by country · publication coverage, not incident prevalence

This view mixes multiple possible relations—actor origin, target geography, infrastructure, policy context, and publication scope. A follow-on incident dataset would need explicit relation types before geographic risk scoring.

What defenders should do with these findings

The corpus supports a practical defensive program, not an “AI attack” product category.

  1. Prioritize identity and social-engineering controls. Validate email authentication, high-risk sign-in detection, MFA reset procedures, executive impersonation response, help-desk verification, and vishing escalation.
  2. Keep browser and endpoint telemetry connected. Fake AI sites, generated scripts, malware delivery, browser abuse, and token theft cross those boundaries.
  3. Instrument code and vulnerability workflows. Monitor abnormal repository access, secret exposure, package publishing, CI/CD identity, exploit testing, and AI-assisted code use according to policy.
  4. Measure behaviors, not AI branding. Detections should target credential access, execution, persistence, exfiltration, evasion, C2, and fraud evidence. “Generated by AI” is often unobservable or irrelevant to containment.
  5. Create evidence-level provider-abuse records. Separate provider mention, confirmed account abuse, model-output evidence, API telemetry, and publisher inference.
  6. Preserve negative evidence. If a report forecasts a technique but provides no incident, telemetry, sample, or IOC, record that limitation.
  7. Build an incident-normalization layer before prevalence claims. Cluster publications into campaigns/incidents, assign relation types, record confidence, and select one canonical incident row before calculating attack shares.
FieldRequired question
Incident identityDoes this publication describe a unique event or repeat another source?
AI roleWas AI observed, source-reported, inferred, demonstrated, or forecast?
ActorIs attribution explicit, current, and supported?
Provider/modelWas use confirmed, merely mentioned, or reported by the provider?
BehaviorWhat exact action occurred, and what telemetry supports it?
OutcomeDid the action succeed, fail, remain experimental, or remain unknown?
Victim/sector/countryWhat relation does the value have to the incident?
MetricWhat is the denominator, unit, time window, and source methodology?
IOCIs the observable exact, active, scoped, and safe to operationalize?

Explore the connected research ecosystem

This study is the statistical layer of a broader evidence-to-action workflow on 1200km. Use the connected modules according to the question you are trying to answer:

Limitations and reproducibility

Sampling limitations

Measurement limitations

Reproduce the artifacts

This publication exposes the normalized dataset snapshot and chart inputs, but it intentionally does not republish the downloaded third-party HTML and PDF archive, copied evidence excerpts, or local archive paths. Source publishers retain rights in their original reports. The working research project regenerated the normalized data from that private evidence archive and then generated the article, dashboard, and visualizations. Dependency versions were not pinned in the original collection workflow, so exact byte-for-byte rebuilds require the preserved working environment; the public files support analytical reproduction of the reported counts and charts.

Primary analysis files:

Conclusion

The strongest conclusion from this corpus is not that AI has replaced established attacker tradecraft. It is that AI is repeatedly described as a flexible accelerator across identity abuse, reconnaissance, evasion, malware work, localization, vulnerability research, influence operations, and fraud. The literature spans the full intrusion lifecycle, but its evidence ranges from incident response and in-the-wild observation to forecasts, underground claims, and controlled demonstrations.

That distinction determines how the statistics should be used. Publication counts can show research attention, evidence concentration, and analytical gaps. They cannot produce global attack prevalence, provider market share, unique victim counts, or attribution certainty. The next research milestone is an analyst-reviewed incident layer that clusters repeated reporting, encodes relations, validates evidence, and separates observed operations from experiments and forecasts.

Until then, the dataset is most valuable as a reproducible CTI map: it tells defenders where to inspect, what evidence to demand, and which claims remain unproven.

References

Dataset and methodology

Selected primary research in the corpus

Follow My Work

I publish practical cybersecurity research, CTI workflows, detection engineering notes, malware-analysis projects, AI-security research, open-source tools, labs, and technical guides.