Using Satellites to Find Human Trafficking in Fleets at Sea
Testing a forced-labor detection model on high-risk fishing vessels identified after it was built
Interactive: where high-risk fishing fleets came into port in 2024. Open the port explorer in a full window, or read how it works.
In 2021, a team of researchers published a paper in PNAS with a striking claim: satellite data on how fishing vessels move can reveal which ones are likely using forced labor. The paper, “Satellites can reveal global extent of forced labor in the world’s fishing fleet” by McDonald et al., trained a machine-learning model on vessels with reported labor abuses and applied it to about 16,000 vessels.1 Between 14% and 26% of them came out as high risk.
The authors found 193 reported cases of forced labor, but only a fraction of those vessels had enough satellite data to use, and the baseline model trained on 21 of them. A critical reply in the same journal (Swartz et al., 2021)2 raised concerns about the small dataset, the model’s assumptions, and how it was validated, pointing out that the same 21 cases were used both to train the model and to evaluate it. (The authors responded in a reply of their own3).
Forced labor at sea is one of the hardest human trafficking-linked abuses to detect. Workers are recruited with false promises, charged fees they can’t repay, and put on vessels that don’t return to port for months. If satellites could point inspectors to the vessels where this happens, it would change how we address it.
Five years later, and we can test whether they can with the new cases since then. Since the paper was written, journalists and U.S. Customs have identified dozens of fishing vessels with forced labor on board. The model never saw them. So I asked a simple question: did it flag them before anyone else did?
My answer: not in the way the paper suggests, but yes in a way that matters more for policy. The model can tell you which fleets to worry about. It can’t tell you which boats in those fleets are the problem. That distinction should inform how governments use this kind of tool.
How the model works
The model uses Automatic Identification System (AIS) data, the position signals that most large vessels broadcast, processed by Global Fishing Watch4. For each vessel and year, it calculates features like how far the vessel went from port, how many hours it fished on the high seas, how often its signal went dark for more than 24 hours, how many port visits it made, and how often it loitered at sea. It also includes vessel traits like engine power, gear type, and flag.
The model is a random forest trained with positive-unlabeled (PU) learning, a method for data where you know some positives but can’t be sure the rest are negative. Most vessels without a report aren’t confirmed clean and are just unlabeled. To handle this, the model is trained 100 times, each time on all 21 known cases plus an equal-sized random sample of unlabeled vessels, and the 100 results are averaged. The output is a risk score from 0 to 1 for every vessel in every year from 2012 to 2018.
The new cases
I collected vessels identified after the model’s training data from two sources:
- U.S. Customs and Border Protection, which has issued Withhold Release Orders against individual fishing vessels, and one against an entire fleet, since 2020.5 These orders block imports and rest on a government investigation.
- The Outlaw Ocean Project, whose investigation of China’s distant-water fleet tagged 35 vessels with documented forced labor in its public vessel database.6
I matched each vessel to its identity in Global Fishing Watch, using IMO numbers where possible, and checked every one against the original paper’s case database. One vessel appeared in both sources, so I counted it once. I removed one vessel that was already in the paper’s data, one Customs case that matched two different vessels with the same name, one vessel I couldn’t find in the satellite records, and the fleet-wide order, which names a company rather than specific vessels. That left 36 new cases. Sixteen of them had data in 2018, the last year the model scored.
These 16 vessels are the test: 13 squid jiggers and 3 longliners, flagged to China, Taiwan, Fiji, and Vanuatu.
Rebuilding the model
The authors published their code and data7, which is what made this possible, and kept their vessel-level scores anonymous, deliberately, to protect vessels from false accusations, so I didn’t try to reverse that in this analysis. I retrained the model from their public training data using the exact settings in their code: the same 66,368 vessel-years, the same 21 positive cases, the same preprocessing, and the same random forest with 100 bags and 1,000 trees.
The rebuilt model tracks the original closely. For the known forced-labor vessels, where the published file keeps a non-anonymous ID, my scores correlate with the authors’ at 0.88. The pattern across gear types matches, with squid jiggers scoring the highest, longliners in the middle, and trawlers scoring the lowest. It isn’t identical since random forests trained on 21 positives vary from run to run, so I reran every test with five different random seeds and the results moved by less than 0.02.
The model flagged every squid-jigger case, and also most squid jiggers
In my rebuild, all 13 new squid-jigger cases scored above the paper’s high-risk threshold from 2018. On its face, that’s a perfect record. However, look at the base rate; even in the authors’ own published predictions, 75% of all squid jiggers in 2018 are high risk. A model that flags three out of four vessels in a fleet will catch most cases in that fleet almost by default.
Grey bars: McDonald et al. (2021) published predictions, 2018. Red points: my rebuild of the same model, at the paper’s threshold (0.57). No later-identified trawlers had 2018 data.
The next question is whether the model found a way to tell the real cases apart from all the other vessels it flagged. Among vessels it already considers risky, can it rank the ones that were later found above the rest?
Across fleets, this model beats simple rules
Compared with every other vessel in 2018 (12,784 vessel-years), the model ranks the new cases very well. Its AUC, the probability that a randomly chosen case scores higher than a randomly chosen other vessel, is 0.89.
That’s better than simple inspection rules built from the same training data. (A rule that uses only gear type gets 0.75, and a rule that uses gear type and flag gets 0.80).
The model’s advantage over both is clear: 0.14 over gear alone (95% confidence interval 0.09 to 0.18) and 0.09 over gear and flag (0.04 to 0.13). It is able to separate high-risk groups of vessels from the rest better than a blunt rule would.
Inside a fleet, the signal disappears
The harder test compares each case only with vessels of the same gear and the same flag, so Chinese squid jiggers against other Chinese squid jiggers. Here the model does no better than chance.
The combined AUC is 0.48 (95% confidence interval 0.34 to 0.62). In the largest group, 13 cases among 506 Chinese squid jiggers, it’s 0.53. The median score for the cases was 0.672. For the other vessels, 0.668.
Scores from my rebuild of McDonald et al. (2021). Median score: cases 0.672, rest of fleet 0.668. Within-fleet AUC = 0.53.
This matters because of how the cases were found. Most came from one investigation of one country’s fleet, so any model that rates that fleet highly will look good against the whole world. The same-fleet comparison removes most of that effect, and when it does, no detectable signal remains.
Some case vessels were tagged by investigators for AIS transmission gaps, a behavior the model also uses. If investigators and the model were looking at the same signal, the test could flatter the model. Removing those vessels didn’t change the within-gear result.
The model does push a small number of vessels into a low-risk tail, and none of the cases are there. But inside the main cluster, where almost all of the fleet sits, the cases are spread across it like everyone else.
Another way to see this is to imagine an inspector working through the 519 Chinese squid jiggers in order of risk score, highest first. After checking the top quarter of the fleet, they would have found 3 of the 13 later cases. Checking vessels in random order would find about the same number (3.2).
Scores from my rebuild of McDonald et al. (2021). Cases: U.S. CBP and the Outlaw Ocean Project.
Later data points the same way
To use more of the new cases, I rebuilt the model’s inputs for 2019-2022 from Global Fishing Watch’s public data and applied the unchanged model. When I computed the same inputs for 2018, the resulting scores correlated at 0.92 with scores from the authors’ original data. Of the 36 new cases, 28 had data in those years and used one of the three gear types the model covers.
The public data doesn’t include vessel characteristics like engine power and crew size for the newest vessels, so I gave every vessel the typical values for its gear and flag. Within the same gear and flag, the model does slightly better than chance: an AUC of 0.61, with a 95% interval of 0.50 to 0.71.
For the 16 vessels that also appear in the 2018 test, it’s 0.59, and the interval includes chance. Removing the cases that investigators flagged for AIS gaps brings it to 0.58, again with an interval that includes chance. So the later data points the same way as the 2018 test: at best, a weak signal about individual boats.
Vessels that change their identity
While matching the cases to satellite records, I realized how slippery vessel identity is. One vessel sailed under two different names on the same IMO number. Another was renamed in 2023. Six vessels from the same numbered series moved to the Kenyan flag within weeks of each other in mid-2020. One IMO number appeared on more than 15 different transponder IDs, some of them registered to other countries. Changing names, flags, and transponder numbers is a known way to avoid scrutiny and it’s also part of why vessel-level detection from satellite data is hard given the model scores a transponder ID, not a ship.
What this means for forced-labor prevention policy
The debate about satellite detection of forced labor has mostly been about whether the models are accurate. But, what kind of decision is the model accurate enough to support? With the evidence analyzed here, I argue the answer is decisions about fleets and ports, not decisions about boats. I think that points to four changes.
- Use satellites to decide where inspectors go, but not to accuse ships. The model separates high-risk fleets from the rest better than simple rules do and the traffic is concentrated. In 2018, ports in five countries and territories received 54% of all port visits by vessels the model rated high risk. I want to be very clear that we should not use a vessel’s risk score to accuse that vessel of forced labor.
Data: McDonald et al. (2021), PNAS, published port-visit table. Visits with unknown port country excluded.
The map below applies the model specifically to fleets. Each 2024 port visit counts with the share of its fleet that the model rated high risk, and the default view shows only foreign-flagged vessels, because those are the ones a port state’s inspectors can act on. In 2024, five ports received 42% of these risk-weighted visits, led by Port Louis in Mauritius.
Open the port explorer in a full window
Inside a flagged fleet, I would recommend inspecting at random. Ranking vessels by risk score finds cases no faster than chance, as the inspector chart shows. Random selection within high-risk fleets finds cases just as fast, is harder for operators to predict, and treats every vessel in the fleet the same way.
Treat identity changes as a warning sign. Several case vessels changed names, flags, or transponder numbers, and six moved to the same new flag within weeks. A model that scores transponder IDs loses a vessel every time it does this. Requiring permanent IMO numbers and flagging sudden mass reflagging would make both human and satellite monitoring harder to escape.
Limits
Sixteen cases in 2018 and 28 in later years are enough to show the pattern above, but not enough to rule out a modest vessel-level signal. The confidence intervals are wide.
The cases are not a random sample. Most come from one investigation of one country’s squid fleet, so other fleets could look different. Six of the new cases, the vessels that moved to the Kenyan flag, are purse seiners, which the model doesn’t cover, so they aren’t in either test.
The comparison vessels are unlabeled, not confirmed clean. Some of them probably also use forced labor, which pulls the measured AUC down. The true within-fleet performance is likely somewhat higher than what I measured, though I have no way to know by how much.
My model is a close rebuild of the original, not the original itself. For the later years, the inputs are rebuilt from public data and match the authors’ closely but not exactly. And because the comparison vessels all come from the 2018 fleet, the later test can’t compare the newest case vessels with other new vessels.
The bottom line
Satellite detection of forced labor was pitched as a way to find the boats. I don’t think the data shows it does that, but it tells us which fleets to watch and which ports they use. The most effective use of this is a map for where inspectors should go, paired with random inspections.
If fleet-level targeting is where the value is, the natural next question is whether the authorities that act on fleets are targeting the right ones. U.S. Customs has already issued one order against an entire fishing fleet. In the next post, I’ll look at whether its enforcement follows the goods most at risk of forced labor.
Methods and code: https://github.com/aadhavr/fleets-not-boats
The code downloads the authors’ original data and applies all changes locally; it doesn’t redistribute modified versions of their files. Data sources: McDonald et al. (2021), PNAS; Global Fishing Watch; U.S. Customs and Border Protection; the Outlaw Ocean Project.
Footnotes
McDonald, G. G., et al. (2021). “Satellites can reveal global extent of forced labor in the world’s fishing fleet.” PNAS 118(3). https://www.pnas.org/doi/10.1073/pnas.2016238117↩︎
Swartz, W., et al. (2021). Reply in PNAS 118(19), e2100341118. https://www.pnas.org/doi/10.1073/pnas.2100341118↩︎
McDonald, G. G., et al. (2021). “Reply to Swartz et al.” PNAS 118(19), e2104563118. https://www.pnas.org/doi/10.1073/pnas.2104563118↩︎
Global Fishing Watch. https://globalfishingwatch.org↩︎
U.S. Customs and Border Protection, Withhold Release Orders and Findings. https://www.cbp.gov/trade/programs-administration/forced-labor/withhold-release-orders-and-findings; fleet-wide order against Dalian Ocean Fishing: https://www.cbp.gov/newsroom/national-media-release/cbp-issues-withhold-release-order-chinese-fishing-fleet↩︎
The Outlaw Ocean Project, Bait-to-Plate vessel database. https://b2p.theoutlawocean.com/vessels↩︎
Authors’ code and data. https://github.com/emlab-ucsb/slavery-in-fisheries↩︎