↓ Skip to main content
Large Scale Threat Actor Attribution from Infostealer Logs

Large Scale Threat Actor Attribution from Infostealer Logs

·1905 words·9 mins

Infostealer logs are usually viewed from the perspective of the victim: stolen credentials, session cookies, autofill data and other information that can later be used for account takeover, fraud or further compromise. But the collection process is indiscriminate. People involved in cybercrime use the same browsers, operating systems and online services as everyone else, and some of them also end up infected. According to latest research by Flare, up to 5% of these victims are actually threat actors themselves.

That creates an interesting attribution opportunity. If an infostealer log contains an authenticated session for a criminal platform, and the same log also contains enough reliable information about the infected user, it may be possible to connect an underground account to a real-world identity. The process could be automated to work across a larger number of cases rather than as a one-off investigation.

This research was also presented at BSides Tallinn 2026 on 25 September 2026 under the title “Identifying 100 Cybercriminals In 1 Hour.” The presentation covered the same methodology described here: selecting a suitable criminal ecosystem, mapping stolen session data to forum accounts, building structured threat actor profiles, extracting higher-confidence identity information from the corresponding stealer logs, and enriching those entities against existing intelligence sources.

Why infostealer data is useful for attribution
#

Infostealers collect considerably more than passwords. Depending on the malware family and configuration, a log can contain device information, saved credentials, autofill data, browser cookies and other browser artefacts. The stolen data is then distributed through Telegram channels, forums and other parts of the stealer-log ecosystem.

Alt text
Example of infostealer collected data. Source HudsonRock.

Most analysis of this data naturally focuses on the people whose accounts are later abused. For this project, we were interested in the smaller subset of victims who were themselves participating in criminal communities.

Recent research indicates that around one in twenty stealer-log victims may be threat actors, or roughly 5%. That number should not be treated as a precise measurement of every stealer dataset. Logs frequently contain duplicates, the same machine can be infected multiple times, and sellers may inflate the size of their collections. But even with those limitations, the overlap is large enough to be useful. If a threat actor’s own machine appears in a stealer dataset, the log may contain both traces of their underground activity and information that helps identify the person using the machine. The challenge is linking those two sides with enough confidence.

Choosing the target
#

Alt text
Example list of candidates.

We started by looking at several established cybercriminal forums and related ecosystems. Patched, Hack Forums, Altenen and others all appeared with significant volume in the stealer data we were working with. In the comparison shown during the presentation, some of these platforms appeared in thousands of records, with Altenen appearing in 10,336.

Alt text
Candidates with infostealer log volume.

High volume alone was not enough to make a platform useful for this approach. We needed a reliable way to connect a stealer log to a specific account on the platform.

The requirements were therefore fairly practical. The platform had to be clearly associated with criminal activity, the stealer log needed to contain a useful entry point, and that entry point needed to allow us to identify the corresponding platform account. The presentation described this distinction as the difference between weaker signals such as URL or credential presence and stronger signals such as an authenticated session cookie.

A username and password for a forum can be useful, but it does not necessarily tell us much on its own. Credentials can be old, reused or shared. A session cookie provides a stronger link because it represents a browser that was authenticated to the service.

The next question was whether the cookie could be mapped back to a user.

This varied by platform. On some MyBB-based sites, the relevant cookie values were encrypted or otherwise difficult to translate directly into an account identifier. Altenen was more convenient because it uses XenForo and, in the logs we examined, the XenForo session data allowed us to recover the corresponding user ID relatively easily.

Alt text
Altenens.is xenforo session cookie.

That gave us a clean path from an infostealer log to an underground account.

Mapping a stealer log to a threat actor account
#

The first half of the pipeline focused on the criminal persona:

infostealer log → session cookie → user ID → forum account → threat actor card

For logs containing cookies belonging to the target platform, we extracted the relevant session value and then recovered the user ID encoded or associated with it. The presentation walks through this sequence directly: first extract the session cookie, then derive the user ID.

Once the account ID was known, our browser automation was used to visit the user’s profile and collect the information associated with it. This included the profile itself, account statistics and the user’s posts.

Alt text
Wayback view segment of the scraped profiles mapped from infostealer session cookies to users on the platform.

For each mapped account, the pipeline collected the profile, found the user’s postings, calculated basic statistics and summarized the topics discussed by the account. The resulting structure included the platform username, post and thread counts, active years and categories of interest. We referred to this normalized result as a threat actor card.

Alt text
Example threat actor card.

Identifying the person behind the infected machine
#

The second half of the workflow used the same stealer log, but for a different purpose.

Once we had mapped the browser session to an underground account, we examined the rest of the log for information that could help identify the user of the infected machine.

This part required more caution as a typical stealer log can contain large numbers of usernames, email addresses, credentials and URLs, and it would be wrong to assume that all of them belong to the victim. A browser may contain credentials for work systems, customer accounts, family members or services used only occasionally. Particularily, a infostealer victim operational in cybercrime domain has a high chance of mingling with stolen credential and identities, contaminating the log even further.

The pipeline therefore extracted entities from the log and assigned confidence to them. Only higher-confidence entities were then enriched against our intelligence sources.

The enrichment step was important because a single identifier is often not enough. A username may be reused by several people. An email address may be old. A phone number may appear in unrelated datasets. The goal was to look for combinations of identifiers that supported each other. Hence the pipeline then combined the raw data and the enrichment results into an entity graph while retaining confidence on the relationships.

The final normalized identity card could contain fields such as the person’s name, email address, nationality and phone number when those values were supported by the available evidence.

Alt text
Example real identity card.

Putting both sides together
#

At this stage, the same stealer log had been processed in two directions.

Alt text
Example output of the flow.

On one side, the stolen session data identified an account on a criminal platform. The account’s profile and posting history were then collected and normalized into a threat actor card.

On the other side, the personal information in the same log was extracted, scored, enriched and normalized into a real-world identity card.

Alt text
Overall flow.

Why this is relatively low-hanging fruit
#

Most of the required information has already been collected by the infostealer.

The session cookie provides a way into the criminal-platform side of the attribution problem. The rest of the log provides candidate identifiers for the person using the infected machine. The remaining work is largely around parsing, selecting useful signals, enrichment and normalization - Public account information could be collected after the account had been identified, while the attribution data came from the existing stealer logs and our intelligence sources.

That makes the method well suited for automation. Once the extraction logic for a particular platform has been implemented, the same sequence can be repeated for each matching stealer log with relatively little manual work. There are still platform-specific differences - The XenForo cookie structure made Altenen particularly convenient for this proof of concept, while other forum software could be more tedious to map. The method is therefore not automatically portable to every criminal community, but the broader workflow remains the same wherever a strong link between a stolen session and a platform account can be established.

Additionally, the infostealer log file provides such a high volume of data, that even if a threat actor decides to reinvent their whole underground identity, there is still high chance of contamination allowing to link the new alias back to the old alias.

Results
#

The full proof of concept produced 546 alias and identity pairs during approximately three and a half hours of processing. The presentation recorded an average processing time of around 20 seconds from a raw log to the final pair once the pipeline was running.

The research provided another useful observation: 141 of the 546 aliases, approximately 25%, were active across multiple criminal platforms.

Alt text
Threat actor distribution across multiple criminal platforms.

Limitations
#

Automating the process does not remove the uncertainty that comes with attribution.

A session cookie is a relatively strong signal, but it still describes a session on a machine rather than proving that one particular person was physically using the computer at that moment. Machines can be shared, browser profiles can contain multiple users’ data, and sessions can sometimes be transferred between systems.

The identity side has similar limitations. A stealer log can contain credentials and personal information belonging to several people, which is why entity confidence and enrichment are important parts of the workflow. A single email address, phone number or username should not automatically be treated as the identity of the machine owner. For the output to be actionable, it still requires human validation.

The source datasets also need to be handled carefully. Multiple infections may refer to the same machine, datasets may overlap, and advertised log volumes are not necessarily equivalent to unique victims.

For these reasons, the pipeline retained confidence information instead of reducing every result to a binary identified/not-identified decision. The stronger cases are those where several independent pieces of information converge and where the link from the infected browser to the criminal account is clear.

Conclusion
#

Infostealer data contains more than credentials that can be used against ordinary victims. A portion of the infected machines also belong to people participating in the criminal ecosystems where those same logs are traded and consumed.

For this project, the useful combination was an authenticated criminal-platform session together with sufficiently strong personal identifiers inside the same stealer log. Where the forum’s session format allowed the account to be recovered reliably, the rest of the process could be automated: collect the account’s public activity, normalize it into a threat actor profile, extract higher-confidence identity information from the corresponding stealer log, enrich those entities and build a structured attribution record.

The proof of concept produced 546 alias and identity pairs within 3 hours and 30 minutes of runtime, and showed that much of this work can be handled as a repeatable data-processing pipeline rather than a separate manual investigation for every account. The approach will not work equally well across every forum or every infection, but where the underlying signals are strong enough, infostealer logs can provide another useful source for threat actor attribution.

Related