South Korea’s Ministry of Science and ICT has selected a 33-organisation consortium, led by Naver Cloud, to develop a 700 billion parameter security-focused AI foundation model.South Korea’s Ministry of Science and ICT has selected a 33-organisation consortium, led by Naver Cloud, to develop a 700-billion-parameter security-focused AI foundation model. S2W, whose DarkBERT model was trained on over 6 million dark web pages, has joined the consortium to contribute hidden-channel data that the surface web can’t provide. The two resulting models, built on Naver Cloud’s HyperCLOVA X and LG AI Research’s EXAONE, will be tested at critical national facilities including aerospace, nuclear power, and financial institutions, then released as open source for commercial use.
The project was officially awarded on September 3, 2026. On September 14, S2W, the Korean threat intelligence company that built DarkBERT, the first major AI model trained exclusively on dark web content, confirmed it has joined the consortium as the primary contributor of hidden-channel intelligence data.
The goal is a 700-billion-parameter AI that understands not just cybersecurity in theory but the actual language, operational patterns, and evolving tactics that criminal threat actors use in underground forums and marketplaces that never appear on the open internet.
No country has announced a government-funded security AI at this scale trained on this kind of data before.
What S2W Actually Does And Why Dark Web Data Changes Everything
S2W was founded in September 2018 in South Korea and built its entire business around a single premise: the most valuable early signals about emerging cyber threats don’t appear on the surface web. They appear in underground forums, dark web markets, private Telegram channels, and criminal communities where threat actors discuss operations before they’re launched, trade stolen credentials before they’re deployed, and sell access to breached networks before the victim knows anything happened.
That intelligence gap, between what’s visible on the open internet and what’s actually happening in criminal communities, is precisely what S2W specialises in closing. Its multidomain cross-analysis technology crawls and processes this hidden-channel content continuously, feeding AI models that translate raw criminal activity into actionable threat intelligence.
This is the same intelligence gap we’ve covered extensively in tracking how supply chain attacks show up on the dark web before they become public incidents. By the time a breach makes the news, underground forums have often been discussing the target, the access methods, and the early sale of stolen credentials for weeks. S2W is building AI to read those signals systematically.
DarkBERT: The Foundation Under the Foundation Model
To understand what S2W is bringing to this project, you need to know what DarkBERT is.
In May 2023, S2W and KAIST published a research paper that was accepted at ACL 2023, the Association for Computational Linguistics conference, which is among the most competitive venues in natural language processing globally. The paper introduced DarkBERT, a language model trained exclusively on dark web content.
DarkBERT was built on top of RoBERTa, an architecture developed by Facebook AI Research in 2019 and one of the most robust encoder-based language models available. S2W fed it over six million pages of dark web content, crawled using Tor, and trained it over 16 days across two datasets. Some categories were filtered before training, illegal images, victim organisation names from leak posts, and explicit threat content directed at specific individuals. Everything else went in.
The result was a model that does something general AI models genuinely struggle with: it understands criminal vocabulary. Dark web forums use rapidly evolving slang, obfuscated language, intentional misspellings, abbreviations that exist nowhere else on the internet, and operational codes that shift constantly to evade detection. A model trained on Reddit and Wikipedia reads those forums the way a monolingual English speaker reads a text in a different dialect. They can parse individual words but miss the meaning entirely.
DarkBERT clocked 90% accuracy in classifying illegal activities on dark web pages, according to S2W’s reporting. More importantly, it can detect when confidential data appears in a leak post and infer keywords associated with emerging threats, capabilities that matter enormously for early warning intelligence and that general-purpose models can’t match.
DarkBERT was never made publicly available. Requests for academic access could be submitted to S2W, but the model stayed out of commercial circulation. The reason is straightforward: a model trained on criminal content that can understand and reason about that content is a dual-use tool. What detects criminal activity can also assist it.
The 700B government project is DarkBERT’s next chapter, with the hidden-channel data S2W has accumulated since 2023 โ continuously updated, feeding a vastly larger model built for national security deployment.
The 700B Model: What’s Actually Being Built
, The scale here is worth pausing on. 700 billion parameters puts this model in territory occupied by only a handful of systems globally. GPT-4 is estimated to have around 1.8 trillion parameters across its mixture-of-experts configuration, but most security-focused AI models that exist today are substantially smalleroften in the 7B to 70B range.
The chosen architecture is Mixture-of-Experts (MoE). Rather than activating all parameters for every query, MoE models route each input to the most relevant subset of specialised sub-networks, “experts”, which makes the system computationally efficient while preserving the depth of a much larger model. It’s the same architectural approach used by several leading frontier models including OpenAI’s GPT-4.
Critically, the consortium is building not one model but two distinct 700B-class systems. The first, built on Naver Cloud’s HyperCLOVA X foundation, focuses on defensive security capabilities, detecting threats, classifying attacks, and supporting incident response. The second, built on LG AI Research’s EXAONE, targets attack response โ understanding adversary tactics well enough to predict and counter them.
These two models will operate together through what the consortium calls a “harness” structure (ํ๋ค์ค), a coordination layer that connects and controls multiple AI security agents simultaneously. The aim is to create a system that operates coherently across different security functions rather than requiring a human operator to manually synthesize outputs from separate tools.
Hardware deployment is already underway without waiting for the formal project launch. Naver Cloud has pre-deployed 4,000 of its own NVIDIA B200 GPUs for pre-training. The government provides 256 additional GPUs, and LG adds 256 H200 GPUs. That’s a significant private contribution ahead of official funding, which signals how seriously the consortium’s lead organisation is treating this project.
This aggressive hardware pre-positioning has direct parallels to what we’ve seen in the AI capability competition this year, as detailed in our coverage of how Chinese AI labs ran 190 million unauthorised exchanges against Claude to close capability gaps. South Korea, watching that competition from close range, is clearly not content to be a passive participant.
Who Else Is in the Consortium
The 33 organisations in the Naver Cloud-led consortium span industry, academia, critical infrastructure, and the public sector in ways that reflect exactly which sectors the model is designed to protect.
The industry participants include LG CNS (enterprise IT and security services), LG U+ (telecommunications), and S2W. The academic institutions include Seoul National University and KAIST, the same institution that co-developed DarkBERT. The critical infrastructure participants are the most telling: Korea Aerospace Industries (KAI) and Korea Hydro and Nuclear Power (KHNP).
KAI is South Korea’s primary aerospace and defence manufacturer. KHNP operates the country’s nuclear power plants. These aren’t research organisations dabbling in AI. They’re the operators of facilities that represent the most sensitive and highest-value targets for advanced persistent threat (APT) groups. Their presence in the testing phase means the security AI will be validated against exactly the kind of infrastructure that sophisticated state-sponsored attackers have historically targeted.
The Financial Security Institute, which oversees cybersecurity standards for South Korea’s banking system, is also involved. Testing across nuclear power, aerospace, defence, and financial infrastructure is about as comprehensive a real-world validation environment as this kind of AI could get.
SKT had assembled its own competing consortium, which included Upstage, SK Shieldus, AhnLab, Igloo Corporation, RaonSecure, and nine other domestic security companies, alongside Korea University and Soongsil University. That consortium was the other finalist. The Ministry of Science and ICT ultimately selected the Naver Cloud consortium after a final presentation evaluation. SK Telecom confirmed it would continue building its own AI capabilities independently after losing the bid.
Why “Security Sovereignty” Is the Key Phrase
S2W CEO Seo Sang-duk’s quote from the announcement is short but specific: “S2W will contribute to more effectively responding to increasingly sophisticated cyberthreats and strengthening security sovereignty.”
“Security sovereignty” is a deliberate concept in Korean government AI policy. It refers to a country’s capacity to protect its own digital infrastructure using domestically built AI rather than relying on foreign models, with all the dependency, data exposure, and potential access risks that reliance implies.
This concern has become far more pointed in 2026. The Rhysida ransomware group’s attack on Berlin’s government infrastructure demonstrated that even technically sophisticated European governments can have critical data exfiltrated through compromised IT infrastructure before anyone notices. South Korea, which shares a peninsula with one of the world’s most active state-sponsored hacking programmes in North Korea’s Lazarus Group and associated units, has more urgent reason than most countries to develop domestic AI detection capabilities that don’t depend on data routing through foreign servers.
Equally relevant: if a country’s security AI relies on a foreign foundation model, that model’s provider can observe, modify, or withdraw access at any point. Building a domestic 700B security model eliminates that dependency for the applications where it matters most, protecting nuclear infrastructure, aerospace development, and financial systems.
The Open Source Commitment
One aspect of this project that sets it apart from typical government security AI efforts is the planned open source release. The consortium has stated it intends to release the models as open source for commercial use.
Additionally, the project will provide free vulnerability and cyberattack detection checks for small and medium-sized enterprises through the Naver Cloud marketplace. This addresses a real problem: the companies most vulnerable to sophisticated AI-assisted attacks, small businesses without dedicated security teams or budgets, are usually the furthest from access to advanced threat intelligence tools.
A 700B security model that can detect vulnerabilities and analyse threat data, made freely available to SMBs through a cloud marketplace, could meaningfully lower the barrier to meaningful security coverage for businesses that currently operate with little more than basic antivirus protection.
That access gap has real consequences. As our coverage of how dark web carding shops and synthetic identity fraud supply chains operate has shown, small businesses are frequently the weakest link in larger criminal targeting chains, not because their data is inherently valuable, but because their defenses are easily overcome.
Why This Matters Beyond South Korea
The specific architecture choices and deployment targets in this project reflect a broader shift in how governments are thinking about AI and national security.
General-purpose AI models are trained for breadth. They know a little about everything. Security-specialised models trained on domain-specific data, including dark web intelligence, have the potential to outperform general models significantly on the narrow but critical tasks of threat detection, attack classification, and vulnerability identification. The same principle that makes DarkBERT more accurate than GPT-based models at classifying dark web illegal activity applies at scale: specialised training data produces specialised capability.
The 700B scale combined with MoE architecture means this model can maintain that specialisation depth while operating efficiently across multiple parallel security domains, exactly what an organisation monitoring nuclear power, aerospace, finance, and telecommunications simultaneously needs.
For context, consider how threat intelligence currently works at scale. Human analysts monitor underground forums for mentions of specific organisations, new malware families, and credential sales. They can only read so many posts per day. A model trained on dark web data that can process millions of forum posts continuously, flag emerging threats in real time, and automatically correlate criminal activity with known threat actor groups multiplies that analytical capacity by orders of magnitude.
This is the kind of capability that makes the early warning gap that preceded attacks like the Bank of Baroda breach significantly harder for attackers to exploit. The signals were there in underground channels before the breach became public. Reading those signals faster than humans alone can manage is precisely what dark-web-trained AI is designed to do.
What Happens Next
The Naver Cloud consortium is already in pre-training with its 4,000 B200 GPUs deployed. The formal project timeline hasn’t been publicly specified, but the combination of pre-deployment hardware and a 33-organisation consortium with active critical infrastructure partners suggests a development track aimed at real-world deployment within the next 12 to 18 months.
S2W’s specific contribution, continuous dark web data collection and hidden-channel intelligence, is an ongoing operational input rather than a one-time dataset. The model will be updated as the criminal landscape evolves, which is the only approach that keeps a threat intelligence AI from becoming a historical artifact. Criminal forums change terminology, new malware families emerge, and operational codes rotate constantly. A model fed static training data becomes stale within months.
For the organisations watching from outside South Korea, the project is a proof of concept for a model that other governments have discussed but not yet funded at this scale: a sovereign, domain-specific security AI trained on data the open internet never surfaces, deployed against the country’s most sensitive infrastructure.
The question that follows naturally from that concept is where similar efforts are being built elsewhere, and whether they’re being built fast enough.
Frequently Asked Questions
What is DarkBERT?
An AI language model developed by S2W and KAIST in 2023, trained on over 6 million dark web pages crawled using Tor. It can detect confidential data leaks, classify illegal activities with 90% accuracy, and infer cyberthreat-related keywords from underground forum content. It’s built on the RoBERTa architecture and was accepted for presentation at ACL 2023.
What is S2W?
A South Korean cybersecurity company founded in 2018, specialising in threat intelligence collected from dark web and hidden-channel sources. Selected as a World Economic Forum Technology Pioneer in 2023 and listed on KOSDAQ. It’s contributing its dark web data intelligence capabilities to the new 700B security AI project.
What is the 700B model being built?
Two 700B-parameter security-focused AI models using a Mixture-of-Experts architecture, one built on HyperCLOVA X (defensive), one on EXAONE (attack response), developed by a 33-organisation consortium led by Naver Cloud under a South Korean government mandate.
Who else is in the consortium?
LG CNS, LG AI Research, LG U+, KAIST, Seoul National University, Korea Aerospace Industries (KAI), Korea Hydro and Nuclear Power (KHNP), the Financial Security Institute, and 24 other organisations. It will be tested at critical national facilities across aerospace, nuclear, and financial sectors.
Why does dark web data make the model different?
General AI models are trained on surface web content and struggle with criminal community language, slang, and operational patterns. Dark web-trained models understand the actual vocabulary threat actors use in underground forums, which produces more accurate threat classification and earlier detection of criminal activity.
Will the model be publicly available?
Yes. The consortium has committed to releasing the models as open source for commercial use, with free vulnerability scanning for SMBs through the Naver Cloud marketplace.