Ransomware-as-a-service (RaaS) has increased the scale and complexity of ransomware attacks, yet these groups' internal operations are hard to observe because the activity is illegal. The 2022 Conti chat leak, which followed the operator's declaration of support for the Russian government, offers direct material. Applying Natural Language Processing and Latent Dirichlet Allocation to 168,711 leaked Jabber chats from 346 actors, June 2020 to March 2022, the study keeps the 137 actors with enough text. Only 4% held specialised discussions; 96% were all-rounders.
Research Objectives
- Identify the main discussion topics in Conti's leaked internal chats.
- Measure how those topics are distributed across actors to gauge specialisation.
- Assess whether the tech and non-tech mix fits the firm-like view of large RaaS operators.
- Compare topic distributions with roles security researchers assigned to well-known actors.
Methodology
- TheParmak's open GitHub repository of English-translated Conti Jabber logs: 168,711 chats among 346 actors, 21 June 2020 to 2 March 2022, with a gap from 16 November 2020 to 29 January 2021.
- Chats were aggregated per actor, with channel-broadcast duplicates removed as they distorted the corpus.
- NLP handled normalisation, stop word, punctuation and link removal, tokenisation and lemmatisation; actors with under 100 relevant words were dropped, leaving 137.
- LDA models ran in MALLET 2.0.8 via the gensim wrapper; coherence scores, WordClouds and pyLDAvis semantic space selected a five-topic model.
- Topics were named from word lists and actors' discussions, then compared with four qualitative blog analyses.
Key Findings
- Five topics span the corpus: Business, Technical, Internal tasking/Management, Malware, and Customer Service/Problem Solving.
- Non-tech talks averaged 44.2% (SD 21.9) of an actor's corpus, tech talks 31.8% (SD 23.0), and Customer Service/Problem Solving 24% (SD 17.6).
- Internal tasking/Management appeared in almost every actor's corpus. Six of 137 actors (4.38%) had 95% or more of their talk in one topic; the other 95.62% were all-rounders.
- Topic distributions broadly matched researcher-assigned roles: the "Big boss" and "Effective head of office operations" had about 99% of their corpus in Business; testers and coders were Technical-dominant; five managers Internal tasking/Management-dominant.
- Not every case lined up: three actors described as managers of technical teams had Malware dominant, read as hands-on managers, and the sources disagreed among themselves over one actor (79.6% Customer Service/Problem Solving), called both manager and Chief Operating Officer. The authors caution that roles read from conversation may not be perfectly accurate.
Recommendations
- Extend the analysis to the leak's Rocket.Chat logs.
- Repeat the work on the original Russian logs, since slang and abbreviations may be lost in translation.
- Account for corpus size, chat timeline and member status; talkative actors, newcomers, affiliates or customers may differ.
- Combine topic modelling with qualitative thematic analysis, since LDA uses word frequency and cannot capture context.
- Examine other cybercrime organisations; Conti is one of the biggest RaaS operators and may be an outlier.
Key Takeaways
- The five topics point to an enterprise-like organisation where coordination, management and customer service absorb substantial effort; large RaaS operations need a workforce skilled beyond technical abilities.
- Specialisation is confined to a small minority, while most members coordinate the operation across several topics.
- For Thailand and ASEAN, where as-a-service cybercrime hits organisations in both sectors, the study offers a reproducible way to read organisational structure from leaked corpora.
References
Ruellan, E., Paquet-Clouston, M., & Garcia, S. (2024). Conti Inc.: Understanding the internal discussions of a large ransomware-as-a-service operator with machine learning. Crime Science, 13, Article 16. https://doi.org/10.1186/s40163-024-00212-y
Full text (Open Access): https://link.springer.com/article/10.1186/s40163-024-00212-y







