Microsoft says email spammers are adopting ASCII smuggling
Recorded: Sept. 9, 2026, 9 p.m.
| Original | Summarized |
Once popular for attacking AI, ASCII smuggling is embraced by spammers - Ars Technica Skip to content Ars Technica home Sections Forum Subscribe Search AI Biz & IT Cars Culture Gaming Health Policy Science Security Space Tech Feature Reviews AI Biz & IT Cars Culture Gaming Health Policy Science Security Space Tech Forum Subscribe Story text Size Small Width Standard Links Standard * Subscribers only Pin to story Theme HyperLight Day & Night Dark System Search Sign In CAN YOU SEE THIS? Once popular for attacking AI, ASCII smuggling is embraced by spammers A once-overlooked block of unicode that’s invisible to humans is gaining ever wider use. Dan Goodin Sep 4, 2026 1:18 pm | 53 Eyeball with a digital line over it.
Eyeball with a digital line over it.
Text Story text Size Small Width Standard Links Standard * Subscribers only Minimize to nav A clever technique used to hide malicious prompts in attacks on AI agents has been adopted by spammers to evade filters on email platforms that are designed to flag unwanted messages used in mass campaigns. Daily hits on the ASCII smuggling signature, a week before and after onset. Daily hits on the ASCII smuggling signature, a week before and after onset.
Daily Unicode-tag signature hits on finance-themed sender domains, log scale, measured every day from February 9 through June 18, 2026. Daily Unicode-tag signature hits on finance-themed sender domains, log scale, measured every day from February 9 through June 18, 2026.
Spammers are embedding Unicode in an attempt to evade filters that search for text, such as dollar amounts and the words “credit” and “term” that are commonly found in their mass emails. By sprinkling the invisible text into the middle of the word “funding,” for example, filters may read the words “fun” and “ding” instead. The receiver, meanwhile, sees the word “funding.” Credit: Credit:
Using special text to camouflage certain trigger words isn’t new. Spammers have used zero-width spaces and non-breaking spaces for decades to achieve similar results. The characters can thwart searches matching a literal string and alter the byte sequence that regex filters hunt for. The spammers likely adopted the hidden Unicode tags because some spam filters had yet to be programmed to detect them. A bigger likely reason for its use is to counteract the advantages made possible by machine learning (ML) and natural language processing (NL) LLMs for use in spam detection. Dan Goodin Dan Goodin Dan Goodin is Senior Security Editor at Ars Technica, where he oversees coverage of malware, computer espionage, botnets, hardware hacking, encryption, and passwords. In his spare time, he enjoys gardening, cooking, and following the independent music scene. Dan is based in San Francisco. Follow him at here on Mastodon and here on Bluesky. Contact him on Signal at DanArs.82. 53 Comments Staff Picks m marsilies What was the intended use case for this character range? Deprecated (twice) but not forgottenThe Unicode standard defines the binary code points for roughly 150,000 characters found in languages around the world. The standard has the capacity to define more than 1 million characters. Nestled in this vast repertoire is a block of 128 characters that parallel ASCII characters. This range is commonly known as the Tags block. In an early version of the Unicode standard, it was going to be used to create language tags such as “en” and “jp” to signal that a text was written in English or Japanese. All code points in this block were invisible by design. The characters were added to the standard, but the plan to use them to indicate a language was later dropped. With the character block sitting unused, a later Unicode version planned to reuse the abandoned characters to represent countries. For instance, “us” or “jp” might represent the United States and Japan. These tags could then be appended to a generic 🏴flag emoji to automatically convert it to the official US🇺🇲 or Japanese🇯🇵 flags. That plan ultimately foundered as well. Once again, the 128-character block was unceremoniously retired. Riley Goodside, an independent researcher and prompt engineer at Scale AI, is widely acknowledged as the person who discovered that when not accompanied by a 🏴, the tags don’t display at all in most user interfaces but can still be understood as text by some LLMs. Officially they are deprecated for use as a language specifier. They've been repurpose as a modifier for flag emojis so more regions can have flags. 🏴gbwls✦ would produce the Wales flag. This is pretty rare though and not well supported. The only usage specified is for representing the flags of regions, alongside the use of Regional Indicator Symbols for national flags. The tag sequences are derived from ISO 3166-2, but sequences representing other subnational flags (for example US states) are also possible using this mechanism. However, as of Unicode version 12.0 only the three flag sequences listed above are "Recommended for General Interchange" by the Unicode Consortium, meaning they are "most likely to be widely supported across multiple platforms" September 4, 2026 at 6:52 pm Comments Forum view Loading comments... Prev story Next story Most Read 1. 2. 3. 4. 5. Customize Ars Technica has been separating the signal from More Contact Manage Preferences |
ASCII smuggling is a technique involving the use of an overlooked block of Unicode characters to embed malicious information in text in a manner that is invisible to human readers but readable by computers, which has recently been adopted by spammers to evade email filtering systems. This method originated as a way to make prompt injections into large language models more stealthy, where malicious instructions were encoded using specific Unicode tags, such as U+E0041 for "A," allowing the LLM to detect the instructions while remaining hidden from human scrutiny. The fundamental mechanism relies on the fact that this block of 128 Unicode tags mirrors the American Standard Code for Information Interchange, possessing characters that are designed to be invisible to humans, yet they exist at the text-processing level. This property allows attackers to obfuscate keywords—such as financial terms or trigger words—by sprinkling the invisible tag characters into the middle of legitimate words. For example, by inserting an invisible tag into a word like "funding," spammers can cause text filters searching for a literal string, like "funding," to fail, while the intended word remains visible to the recipient. This is analogous to using zero-width spaces or non-breaking spaces, but the use of hidden Unicode tags is likely more effective because existing spam filters may not have been programmed to detect them. The shift in focus for spammers is driven by the advances in machine learning and natural language processing models used in spam detection. The author notes that the primary objective is not just to bypass literal string matches but to counteract the capabilities of ML and NLP models. Since these models often process text by splitting it into tokens or sub-word pieces, inserting an invisible character can alter how the tokenizer interprets the text, potentially leading to rare or unknown sub-tokens or changes in token sequences that bypass detection routines. This means that standard email classifiers may miss these obfuscated attacks unless they employ more advanced techniques, such as optical character recognition or deeper reasoning over the context of the entire message. The underlying Unicode structure provides the necessary framework. The Unicode standard includes a block of 128 characters known as the Tags block, which parallels ASCII. Although initially planned for use as language tags or for representing regional flag modifiers, these uses were ultimately abandoned. However, researchers have found that these tags remain understandable to some LLMs, indicating their deeper utility beyond human perception. While some specific flag modification functions have been deprecated, the concept of using this character set as a covert channel for information remains relevant, even if the practical application has evolved from language specification to a method of evasion. Consequently, filtering systems must evolve to better account for this method of obfuscation to maintain effective security against sophisticated attacks. |