Someone might have noticed that our small Lemmy instance was down for the better part of the last day (30-31 August). I wanted to provide some details about what we were doing during this time, and share my thoughts about how we should respond to the threat posed to Lemmy and the Fediverse in general by the gratuitous posting of CSAM (child sex abuse material) as a form of vandalism.
This is a long post, but there’s a lot to talk about. The tl;dr is that we’re not going away, and we’re confident we can protect ourselves and each other. Read on for the details if you like.
Background
This week saw one of the largest Lemmy instances, lemmy.world, make the difficult decision to take down one of the most popular Lemmy communities, lemmyshitpost, after it was brigaded by bad actors posting a flood of CSAM. For most large online platforms, an attack like this would be interpreted as an obvious act of vandalism and cleaned up by moderators or admins. However, this type of behavior poses a more existential threat to decentralized platforms like Lemmy. Because Lemmy instances federate with one another, flooding a large and popular Lemmy instance with this content causes copies of it to be stored on other Lemmy instances, including many smaller ones that probably wouldn’t be worth targeting directly. Worse, deleting the offending posts doesn’t remove those stored copies. Since this content is illegal merely to possess in many jurisdictions, this exposes the admins of these instances – often a single person or a small tech collective – to a legal risk that’s vastly out of proportion with whatever benefit Lemmy might be providing to their community.
What Are Our Options?
Most instances appear to have responded to this attack by blindly deleting 24 to 48 hours’ worth of images corresponding to the times around the attack. While this more or less ensures legal safety from this particular attack, it does nothing to prevent an identical attack in the future, or a series of well-timed attacks that amount to a denial-of-service attack against image posting on Lemmy, by making it legally risky for anyone to host images.
So, what do we small Lemmy instances do? The proposed responses in the community over the past few days seem to mostly fall into one of these categories:
- Shut down our Lemmy instance immediately and move on to a less dangerous project.
- Remove the offending content in hopes the vandals have learned they weren’t able to destabilize the entire Lemmy ecosystem.
- Obtain anonymous hosting in a jurisdiction with weak laws and/or law enforcement and turn a blind eye.
- Develop changes to Lemmy that make it possible to perma-delete content, and to notify federated instances that they should also delete it.
- Develop moderation tools that can detect this content and remove it quickly, whether it originated from our own instance or from one we federate with.
In my personal opinion, #1 is only useful as a last resort, #2 is willfully ignorant, and #3, though plenty anarchic, is patently unethical. This leaves us with the choice between #4 and #5, and I think the decision between them depends on our view of trust.
To many of us, much of the appeal of Fediverse applications is the decentralization of control in favor of individual communities. The principle is that different communities might have different ideas about what content is acceptable or desirable, that each community is free to act on those ideas independently, and that users are free to ally themselves with whatever communities best fit their values. I view option #4 as a challenge to this principle, because it creates an risk of abuse that can only be mitigated through centralization of control. Let’s say that Lemmy vNext ships a new “This is CSAM” button that will delete an image post from a community, and post a message to all federated communities advising them that the content should be deleted. If communities all trusted each other completely to make correct use of this button, it’s pretty basic human nature that that trust would soon be broken by some bad actor misusing the CSAM button to silence someone in a disagreement. So a web of trust will have to emerge, likely separate from the web of trust already implicit in federation. Most communities will probably do the right thing – they will promise to set aside their ideological differences and to work together to fight a form of content that nearly everyone agrees is harmful. But how long will it be before the vandals responsible for attacking lemmy.world stand up a new instance, or compromise an existing one, just to perform a coordinated CSAM spam attack? At that point, the members of this well-meaning web of trust will likely be forced to defederate from unknown instances – moving to a whitelist approach that effectively removes the anonymous public square from the community.
This leaves approach #5: moderate content locally. If we’re going to distribute trust, we must distribute responsibility, and this means that every instance administrator is going to need sufficient tools to combat CSAM independently. I will be the first to admit that the legal situation in the U.S. and most of Europe doesn’t make this easy: there are no publicly available hash lists of known CSAM content that I’m aware of; even if there were, the vandals would be smart enough not to post files that matched it; and any automated solution with a high false positive rate creates a need for human review, which may not be a legally, ethically or psychologically safe operation.
In other words, if we’re going to continue maintaining this or any other zero-trust public platform, we are going to have to be extremely clear about what content is unacceptable, completely transparent with our community about what we’re doing, and willing to respond appropriately if we are challenged legally.
Where Do We Go From Here?
With all of the above in mind, 4d2 dot org will continue operating lemmy.4d2.org, using a combination of AI and human moderation to ensure that unacceptable content neither originates from nor is federated to our instance. This is similar to the approach already taken by most dedicated image hosting apps.
We’re running lemmy-safety, a really fantastic (and urgent!) modification of horde-safety by dbzer0 (@db0@hachyderm.io), which uses the CLIP ViT-L-14 machine vision model to describe the content of images and videos as a string of text, and then classifies images as potential CSAM based on keyword scores. Images are read from the cloud object store where they reside into the RAM of a machine with a reasonably powerful GPU, and discarded after scoring. This has had a false positive rate of 2.7% when run against all existing images in our store. This tool now runs continuously on the scoring machine, and new image uploads are scored within one minute of upload.
The paths to potentially problematic images are stored in a SQLite database by lemmy-safety. We made some fairly trivial extensions to this database’s schema and wrote some accessory scripts to permit human moderation of flagged images. Images can be reviewed in batches by moderators as time permits. When beginning a moderation session, images are synced from object storage to local RAM; they are deleted when moderation is complete.
Images that are approved have a boolean flipped in the SQLite database to ensure they aren’t presented for review again in the future. Images that are marked as CSAM by human moderators have their file sizes, MD5 and SHA256 hashes logged, and are then permanently deleted from object storage. We also automatically collect any server logs that reference the image’s identifier.
Our existing image store has now been reviewed in its entirety, and we’re confident that there is no CSAM being served by this instance. Going forward, we’ll rely on the process above to keep our community safe. I’m also looking into adapting this approach to scan and review unencrypted media on our (much busier) public Matrix instance.
You Have, Like, 20 Users. Why Take the Risk?
With regard to user count, I’ve watched our Matrix server explode from a dozen user accounts to over a thousand in several months. If Lemmy catches on more broadly, and if instances that were opened as an immediate reaction to the Reddit API apocalypse begin to dwindle, it’s probably reasonable for us to think about how we’d handle that kind of scale.
I think it’s fair to ask: since the AI’s false-negative rate is near zero, why not just accept a 2-3% false-positive rate and delete every image that has a whiff of impropriety about it? I’m concerned about that for a couple of reasons. First, it has the potential to create a pretty broken user experience in some scenarios, since we’re deleting images without deleting the posts/comments that reference them. More importantly to me, after reviewing the content that is incorrectly flagged by this iteration of the AI, much of it consistently falls into one of 3 categories:
- Illustrated sexual depictions of characters who appear to be underage. Most of what gets flagged is lolicon, but there’s a pretty strong bias against manga-style illustrations in general. To say that lolicon is controversial would be an understatement, and I have no personal interest in defending it. However, it is legal in most jurisdictions globally, and it’s not what we’re looking for per se.
- Actual pornographic images depicting adults. Troublingly, the AI shows a strong bias toward identifying photos of femme-presenting people with penises as CSAM, regardless of their apparent age. Nothing in the keyword scoring appears to account for this.
- Anti-pedophilia memes that contain relatively little non-text content alongside text phrases such as “sex offender”, which presumably causes CLIP to hallucinate related keywords that trip the filter.
There are some other false detections that are more amusing than troubling: the AI seems to have consistent moral concerns over depictions of baked beans and the Canadian children’s cartoon character Caillou. But in my view, the 3 categories above mirror problematic human behavior that many of us feral internet creatures are already familiar with: the AI often equates manga-style character art with lolicon, vaguely associates LGBTQ+ people with CSAM, and misunderstands memes with ironic results.
We’re hoping that a hybrid approach will prevent us from creating an AI-manipulated “bubble”, however small, that reinforces these existing human biases, and we’re confident that it’s clear to anyone looking in that we are acting in good faith to prevent the spread of content that does harm. 4d2 dot org has relied on the good will and understanding of (most of) our user community for 20 years, and we will do our best to approach this difficult issue with our characteristic transparency.
