This week a 309 MB text file was uploaded to an underground forum, the file contains 3,919,216 lines of login data and the seller has attached the name AlienTxtBase together with the word ‘fresh’.
The entire news item is as far as it goes; all the factors that would determine whether or not you ought to care are contained in the details which almost no one who was reposting the headline took the time to verify.
Begin with the count since it’s the count that’s carrying out all the work and that’s not right. A count of 3.9 million lines translates to 3.9 million rows in a text file; it doesn’t mean 3.9 million people, 3.9 million accounts, or 3.9 million other items that you’d include in a report. A number of the rows are genuine, but there are many duplicates. A portion of them was stolen back in 2019 and has been circulating ever since. And considering what occurred the last time this particular brand name appeared, some of those rows were likely created on the spot.
The job of working out the ratio actually consists in doing it, since the headline is quite adequate without it.
ULP isn’t a stealer log, and that gap is everything
ULP stands for URL:Login:Password and is a file format; it does not constitute a breach and the two are constantly confused.
The format is there since someone took the raw output of an infostealer and discarded most of it. Whenever a stealer such as Vidar, RedLine or Lumma gets onto a machine, it doesn’t just retrieve saved passwords; it also collects session cookies, authentication tokens, autofill entries, crypto wallet files, browser history, system fingerprints, and often includes a screenshot of what was on the screen at that moment. The original log therefore gives a fairly comprehensive view of one compromised computer.
If you convert it into ULP you should keep three fields per line: site, username, and password; the rest go into the bin.
The extent of damage that a file can cause is determined by the deletion. It is the session cookies and authentication tokens that go past the multi-factor authentication (MFA). When a token is valid, the server knows that the user has already been authenticated and therefore does not display a password request, does not send a push notification, and does not require a code to be typed, since the application assumes that the login has already taken place. Tokens that have been stolen can result in an account being taken over without there being a login event to detect.
A ULP line isn’t capable of carrying out any of those actions. It gives the attacker a password to try; this password might have been rotated, it might be incorrect, or it might trigger an MFA prompt which in fact does fire.
When a listing refers to ‘stealer logs’ but the file in fact is ULP, then either the seller doesn’t know the difference or is relying on you not knowing it; researchers have identified this incorrect labelling as a regular practice in underground markets, in which combolists are advertised as stealer logs in order to take advantage of the higher price.
A ULP dump should be regarded as a credential-stuffing wordlist; that’s how it should be seen.
AlienTxtBase has a history, and it’s not flattering
Since the name on the post has most of the credit associated with it, it’s important to know what the name actually refers to.
In February 2025 the Telegram channel known as ALIEN TXTBASE provided a dataset containing about 23 billion rows of credential data. When Have I Been Pwned took in that dataset it extracted 284 million unique email addresses together with the websites at which the passwords had been entered and the passwords themselves. On the basis of its scale it remains one of the largest collections of credentials ever indexed.
Hudson Rock then took it apart.
What they discovered was not a breach, but a combination of old combolists, cleartext password dumps from previous leaks, and a small number of actual stealer logs which were real though several years out of date. The researchers randomly selected nine email addresses and then checked if the mailboxes had actually existed. Eight of them did not exist; the one that did can be traced to a combolist from 2020 which had been circulating publicly for half a decade.
The malware logs contained within the pile had a structure similar to that of data which other underground groups had previously distributed, indicating that the case involves aggregation rather than a new collection.
The brand that this week’s seller used was AlienTxtBase. AlienTxtBase has never been a compromise event; it is a distribution label which has a well-documented tendency to mix a thin layer of genuine stealer output with a much larger amount of recycled and fabricated filler, and then sets the price of the entire collection on the total number of lines.
“Fresh” is a marketing word
The underground sellers have a limited vocabulary. FRESH. PRIVATE. 2026 LEAK. UNSEEN. NOT PUBLIC.
Everything in it cannot be verified. That’s the dark web equivalent of “artisanal.”
The vocabulary is a result of the way pricing is structured; credential data is assessed according to how fresh and valid it is, since a password obtained three days ago from an active infection is much more valuable than one from 2021. Buyers are unable to check whether the data is fresh before making a payment and sellers can offer it for free. The majority of this market is due to that single flaw.
There is also a funnel operating underneath the surface. Public channels are promoting huge ULP combolists at a low cost or for free. The giveaway shows that the operator has a large volume and helps to build an audience. The genuine product—which are the actual stealer logs still containing the cookies and tokens—is then sold privately at a significantly higher price to a much smaller number of buyers.
This leads to a rather unpleasant conclusion: the most prominent drops are generally the ones with the least value, and the material that is worth worrying about is never posted publicly with its file size indicated in the title.
The recycled junk still works
There are limits to skepticism, and it would be a mistake simply to dismiss all of them.
Because of password reuse, old passwords remain in use for essentially an infinite period of time. When Verizon examined the data on infostealer infections in its 2025 DBIR, it found that in the typical case only 49% of the passwords a user has for various services were different from one another. If you round that figure, it means that half of the passwords that an ordinary person uses are in fact copies of a password they use on another service.
A credential that was obtained from a forum in 2019 is still a gamble against that individual’s inbox, their bank, and their employer’s VPN nowadays. It has never been necessary for the credential to be up to date; what matters is that it is reused.
The subsequent attack is on purpose very dull. The same report showed that credential stuffing accounted for a median 19% of the daily authentication attempts among the SSO providers that Verizon studied. Attackers typically attempt each set of credentials once per account, which ensures that they remain below the lockout thresholds intended to detect brute force attacks. Nothing appears wrong from the defender’s point of view.
The cost does not become apparent until much later. In 22% of all breaches Verizon identified compromised credentials as the method of initial access, and discovered that stolen credentials were involved in 88% of the attacks on simple web applications. According to IBM’s Cost of a Data Breach 2025, breaches based on credentials took longer to detect and contain than any other type, with the timeframe being about 292 days and an average cost of roughly $4.67 million.
At company scale the 2026 DBIR shows an average of 2,362 corporate credentials being breached each month from a single organisation’s email domain. Constella discovered that 78% of companies which had recently been breached had their corporate credentials appearing in stealer logs within six months of the breach.
A fair assessment of this week’s upload comes down to something in the middle. The file as a whole is mainly useless. It’s the part of the supply chain which provides enough raw material so that operators can afford to give away millions of lines as a marketing expense that should concern you.
Six checks before you react to any dump
About as fast as the others. Most of the dumps fail one of the first three, so you are generally able to finish in ten minutes.
Start by looking at the format. A series of flat lines in the form of url:login:password with nothing else indicates a combolist. True stealer logs appear as folder structures, one for each infected machine, including separate files for the cookies, the autofill data, the system information and the installed software. A single text file is considered a derivative product.
First, you should find a victim. A genuine breach is linked to an organisation. You can verify or refute a statement such as “4 million records from [Company]”. Whereas “4 million fresh lines” doesn’t name anyone because there isn’t anyone to name.
See if the domains are consistent. When one compromised service provides credentials for another service. Combolists generate a wide variety of domains that have no connection with one another, since they have been put together from a number of infections and numerous old leaks. This widespread distribution is an aggregation fingerprint.
Check the email addresses to see if they are valid. It was this kind of check that caused the February 2025 dataset to be opened. Since fake entries contain made-up addresses, carrying out MX and deliverability tests on a random sample will quickly show how much of the file consists of unnecessary data.
Check the old dumps. In the case where a reasonable portion of your sample is already present in the collections from 2019 to 2023, the situation is likely one of a repackage. Credential intelligence platforms are able to deal with such cases in bulk, while the Pwned Passwords API works well for small samples.
Last, check for session data. Cookies and tokens are what separate genuinely dangerous from merely irritating. Their absence caps how bad this can get. Their presence raises it sharply, because that’s the material that beats MFA.
I should point out that each figure in this article refers to different things when the same term is used. For example, the term “credential” could mean a password, a token, a cookie, a duplicate, or a login that stopped working two years ago, depending on who is carrying out the counting. Therefore, treat these numbers as indicative rather than as actual calculations.
What to actually do
For each individual, the advice is the same no matter whether or not this specific dump is genuine, which is the reason it’s worthwhile carrying out. You should check your exposure on Have I Been Pwned and change any password that you’ve used somewhere else. Enable multi-factor authentication and choose an authenticator app or a hardware key rather than SMS. After that, go to your account settings and sign out of the active sessions on any devices that you don’t recognize, since changing your password doesn’t always end a session that someone else has already obtained.
What security teams should do is regard exposure of stealer logs as an early warning rather than as an incident report. The fact that corporate credentials appear in a combolist does not mean that anyone has gained access to systems; it only shows that an employee’s device was at some point infected, and that is a more significant point since it points in a different direction.
The next steps are affected by this. While rotating the exposed password is straightforward, the important questions to ask are which device generated it, whether that device is still in use within the fleet, and whether the tokens from it have ever been revoked. You should monitor your own domains in the stealer-log feeds. Session revocation should be paired with password resets by default, and valid logins from devices or locations you’re not familiar with should be regarded as suspicious even if the credentials are correct, since in the case of token theft they always will be.
The number was never the story
A text file of 309 MB appeared on a forum bearing a name that is familiar. What it confirms is more useful than the information it contains. Nowadays there is so much infostealer output that operators hand out millions of lines of it as advertising. Recycled passwords continue to work since people keep reusing them. And from the format of the dump you can assess the severity in just thirty seconds whereas the file size will never do so.
When you come across one of these again in your feed, don’t look at the number. Instead, check the format, see if there’s a named victim, and look for session data. If you do those three things and spend a few minutes on it, you’ll be able to tell whether it’s a threat or an ad.
Usually it’s an ad.