In the last few months, I’ve been working on a project called “Sentra”, a system for detecting terrorism-related communication inside encrypted messaging platforms, trying to balance privacy, investigative needs, and practical constraints as much as possible.
In this post, I don’t want to focus on the system’s performance (which, compared to a traditional LLM, achieved an 8x improvement at a lower cost), but on the “philosophy” around such a system and its possible deployment.
I believe that with the evolution of LLMs, they can have a central role as mediators, for example handling tasks where someone wants an outcome without sharing all the underlying information with another human. We often give institutions access to much more information than a decision actually requires. An AI mediator could potentially reduce that exposure, assuming that its access, outputs, and ability to retain information are constrained.
Using a model does not make the question of trust disappear because someone still chooses its training data, writes its instructions, defines what counts as suspicious, and decides what happens after an alert. I don’t think an LLM becomes neutral simply because it has no personal interest in the conversation. The choices of the people building and operating it are still there, but making the system as auditable and transparent as possible can help us examine those choices.
I think that the European debate around what is commonly called “Chat Control” brings many of these questions together. It concerns proposals to combat child sexual abuse online, but it raises a related question, which is under what conditions should private communications become accessible to someone outside the conversation? What’s the limit? Should it be possible in the first place?
My concern is that private conversations could become something inspected by an opaque and inappropriate system in order to establish whether they deserve to remain private or not, and that this process will involve everyone, whether or not they have a background in pedo-pornography or are suspected of it.
I initially approached Sentra through a privacy-security-utility triangle, where you have a three-way tension between each element. If you push on the privacy side, you guarantee full anonymity and untraceability; if you look only at investigative utility, every user should be linked to their passport and every action should be traced; while if you want just usability, every interaction should be as frictionless as possible, requiring no extra steps or barriers and letting the user move through the system quickly and effortlessly. I tried to find the “sweet spot” between those three elements.
Sentra assumes targeted access, only for users on a terrorism watchlist, through a hidden participant or linked device, being aware that this could weaken the trust model of the conversation also for other users. Sentra’s safeguards concern what happens after access has been obtained, they limit the exposure within that process, but they cannot erase the intrusion that makes the process possible.
This is why I think “targeted” has to describe how people enter the system, before their messages are analyzed. Scanning everybody and only showing a small number of conversations to a human still means everybody was scanned. It can also lead to more false-positive messages being read by humans, undermining privacy rights. A system should not be able to use its own suspicion as the only justification for acquiring the information from which that suspicion was produced.
Working on a terrorism detection project is challenging for one main reason: terrorism is a rare event with extremely high consequences. Its rarity also results in a lack of representative data, making it difficult to train and evaluate detection models effectively. For this reason, having real material available and developing a strong synthetic generation process covering different tones, lengths, and writing styles have been crucial to the project.
The conclusion I reached about my own system, is that I don’t think that the math obviously comes out in favor of deployment. What I’m more confident about is narrower, because if a society decides to do targeted monitoring of this kind anyway, then safeguards such as pseudonymization, isolated processing, mandatory human review, and tamper-evident logs could reduce some of its harms, provided that the targeting limits and safeguards actually hold.
The reason I built Sentra and the reason I’m now writing skeptically about it are the same. These systems get built whether or not careful people are in the room, and I’d rather the careful version exist and be argued about openly than leave the design entirely to people who only ever look at one corner of the triangle mentioned at the beginning. The most useful thing a project like this can do might just be to make the trade-off legible enough that the decision to deploy, or not to, gets made with eyes open.
If you’re interested in reading about Sentra, you can find:
- The Sentra paper here (20 pages) which focuses more on the efficiency of the system, reaching more than 8x less expensive and faster performance than using an LLM alone.
- The Sentra extended project here (211 pages), covering the current use of digital tools by terrorist groups, all the details about the dataset creation, benchmarking criteria, and an extensive discussion of the limitations.