Analysis published in the latest issue of the CTC Sentinel by the Combating Terrorism Center warns that current artificial intelligence safety measures may be insufficient to prevent terrorist misuse, with pilot testing showing that many leading AI models continue to provide meaningful assistance when prompted to support terrorist activity despite existing safeguards.
Artificial intelligence has rapidly become one of the defining technologies of the modern era, with new models being released at an unprecedented pace and routinely evaluated for intelligence, performance and cost. Yet, according to analysis in the latest issue of the CTC Sentinel published by the Combating Terrorism Center, one critical question has received comparatively little systematic scrutiny: whether these increasingly capable systems are actually safe when confronted with users seeking assistance for terrorism.
The analysis argues that the most immediate counterterrorism risk posed by AI is not necessarily the emergence of entirely new terrorist capabilities, but rather the willingness of AI models to provide information, guidance and operational support when asked to assist malicious actors. While AI capability continues to advance with each generation of large language models, the study contends that the decisive factor is not what a model knows but what it is prepared to reveal. That distinction, it argues, depends on the effectiveness of the guardrails designed to prevent harmful responses.
To assess that question, Tech Against Terrorism developed its Counter-Terrorism AI Benchmark, described as one of the first systematic efforts to measure how leading AI systems respond to requests linked to terrorist activity. The pilot evaluated nearly 2,500 responses generated by 27 major AI models using a taxonomy covering 152 potential forms of terrorist misuse across multiple operational domains. The models represented developers from the United States, China, France, Canada and the United Arab Emirates, including Claude, GPT, Llama, Qwen, Kimi, GLM, DeepSeek, Falcon, Mistral and Command R.
The findings point to significant weaknesses in current safety measures. According to the benchmark, almost one-third of responses provided meaningful operational uplift that could assist in preparing terrorist activity beyond what an ordinary internet search would readily supply. Researchers emphasise that this figure represents a conservative estimate because every prompt consisted of a single question without follow-up attempts, multi-turn conversations or jailbreak techniques commonly used to bypass model safeguards.
The study also found that the way a request was framed substantially affected the likelihood of receiving assistance. Simply presenting an identical technical request as academic research rather than openly malicious increased compliance from 17 per cent to 42 per cent. According to the analysis, this suggests that many existing guardrails rely heavily on a user’s declared intent rather than reliably evaluating the substance of the request itself.
Equally concerning was the prevalence of what researchers describe as “hedged compliance”. Although many models initially warned users that a request was dangerous or inappropriate, they subsequently supplied the requested information anyway. While overall refusal rates appeared relatively high at 57 per cent, approximately 15 per cent of responses fell into this category, illustrating that warnings alone often failed to prevent potentially harmful assistance.
The benchmark further revealed uneven protection across different aspects of terrorist activity. Requests relating to explosives were refused roughly 80 per cent of the time, but compliance increased substantially when questions focused on operational security, planning and autonomous operations. The analysis concludes that guardrails appear to have been developed around the types of questions designers expected to receive, leaving important operational areas less consistently protected.
One of the report’s most significant findings concerns the distinction between closed and open-weight AI models. Contrary to assumptions that open models necessarily present greater risks, the pilot found examples of both open and closed systems performing comparatively well. However, the study argues that the real strategic concern lies elsewhere: whether safety guardrails can be removed after release.
Researchers describe the practice known as “abliteration”—the deliberate removal of a model’s refusal mechanisms—as one of the most serious emerging challenges. According to the analysis, guardrails can sometimes be stripped through fine-tuning or other technical methods at relatively low cost. Two deliberately de-guardrailed models included in the benchmark complied with 89 per cent and 100 per cent of malicious requests respectively, demonstrating the potential consequences once safety mechanisms are removed.
The report argues that this issue becomes particularly significant for open-weight models, which can be downloaded, modified and redistributed without central oversight. Once publicly released, such models cannot be recalled, updated or withdrawn by their developers. The analysis suggests that developers may therefore be releasing frontier-capability systems whose future behaviour they can no longer control if safeguards prove removable.
The benchmark also challenges the common argument that AI merely reproduces information already available online. Researchers acknowledge that much relevant information exists publicly but contend that AI fundamentally changes how users access, organise and apply that knowledge. During testing, 15 of the 27 models generated at least one chemical, biological, radiological, nuclear or explosive-related response that exceeded publicly available online information, while nine produced technically accurate operational guidance, including detailed procedures involving precise measurements and techniques.
Beyond individual responses, the study argues that AI’s greatest influence may lie in cumulative assistance rather than any single disclosure. Over extended interactions, models may refine ideas, answer technical questions, encourage users or gradually reduce uncertainty, thereby providing incremental support that compounds over time. The analysis also distinguishes between organised terrorist groups and individuals undergoing self-radicalisation, suggesting that the latter may represent the more immediate concern as AI becomes increasingly integrated into online environments where radicalisation already occurs.
According to Tech Against Terrorism’s incident tracker, more than 30 known cases have already involved AI serving as an operational assistant in real terrorist plots or attacks, spanning at least 11 different AI tools and linked to more than 70 deaths. The report argues that this figure is likely to grow as adoption expands and more capable models become widely available.
While acknowledging the considerable benefits of open-weight AI models for innovation, privacy, sovereignty and affordability, the analysis stresses that these advantages should not come at the expense of public safety. Rather than arguing against open releases, it calls for more rigorous engineering of safety mechanisms that cannot easily be removed without fundamentally degrading model capabilities.
The report concludes with several recommendations aimed at strengthening AI safety. These include internationally agreed safety benchmarks that specifically assess terrorist misuse, routine public evaluation of both model capability and disposition, mandatory testing of resistance to guardrail removal before releasing frontier open-weight models, and greater responsibility throughout the AI supply chain, including infrastructure providers, hosting services, model repositories and inference providers.
Ultimately, the analysis argues that while debates surrounding hypothetical existential AI risks continue to dominate funding and policy discussions, more immediate dangers may receive insufficient attention. Technology whose protections can be removed by determined actors shortly after release, it concludes, is insecure by design. As frontier AI systems continue to become more capable and more widely distributed, ensuring that safety measures remain effective may become one of the defining counterterrorism challenges of the AI era.

