Detection must infer a hidden category from operational patterns, not from a direct label.
The DOJ indictment of September 2024 names three Spotify artists: Relaxing White Noise, Meditation Relax Club, and Calmo. Dagens Nyheter's investigation into Johan Röhr's Firefly Entertainment identifies nine pseudonyms publicly, out of a reported 656 or more. But here is the core problem: there is no single Spotify field, no licensing flag, no audit trail that lets you check. Spotify publishes no "ghost artist" label. There is no database you can query. The question forces a methodological pivot. If you cannot specify ghost artists directly, how do you detect them at all? The answer is not to find a single perfect signal, but instead to build seven independent operational measures—each noisy, each individually prone to false positives—and let their convergence on the same artists become the evidence. Convergence itself is the identification move.
The stakes are enforcement, scale, and the broader problem of inferring hidden categories from observable patterns.
Streaming fraud costs the industry billions annually. Bulk-uploaded ghost catalogs dilute royalty pools, diverting payments from legitimate artists. The DOJ prosecution signals enforcement intent, but manual auditing does not scale to millions of artists. What makes this case methodologically important is that the detection framework generalizes. Any latent class—bot accounts on social media, shell companies in supply chains, fraudulent insurance claims—faces the same problem: no single diagnostic, but many noisy signals. Understanding how to infer hidden categories from their observable descendants is a problem that appears across policy, compliance, and investigative work. The stakes are not academic; they are practical and large.
The definitional crisis is why the paper exists. The category must be inferred from production patterns, not from a table cell.
The project opens by asking a deceptively simple question: what does it mean for an artist to be a ghost? Intuitively, you know it when you hear it—bulk-uploaded ambient tracks with no real recording history, no social presence, no verifiable creative identity. But the moment you try to formalize a definition, the boundary dissolves. A legitimate ambient specialist might choose a narrow palette. A ghost might hire session musicians. Spotify has no field for creative authenticity. This definitional crisis is the entire reason the paper exists. You cannot point to a single artist and say "that one's a ghost" by pointing to a table cell. The category must be inferred from the ensemble of production, licensing, and release patterns that cluster together. The paper's claim is that these patterns are real and detectable, even if the label itself is latent.
The arrows mean: if an artist is a ghost, this signal will show a certain pattern. Not the converse.
The causal diagram here is different from what you might expect. In a typical causal-inference problem, you estimate the effect of a treatment on an outcome by closing back-door paths and adjusting for confounders. This is a measurement model. A latent node G—"this artist is a ghost"—sits at the root. Seven observed signals descend from it. Each signal is noisy; each is an imperfect measure of the same hidden truth. The arrows do not mean: if a signal trips, the artist is a ghost. They mean: if an artist is actually a ghost, this signal will tend to show a certain pattern. The threat to identification is not unobserved confounding on a path; it is that the signals are not actually conditionally independent given G. If they measure the same artifact rather than different aspects of the same latent reality, convergence becomes correlation, and correlation is not identification.
The seven signals are audio-feature variance, playlist entropy and release cadence, ISRC registrant concentration, catalog density measured by Herfindahl index on production companies, genre concentration, bipartite neighborhood signal radar combining density with graph degree, and cross-platform discrepancy measured through YouTube views and iTunes presence. Each descends from the ghost-artist hypothesis through a different operational pipeline. Audio variance asks: does this artist's catalog show the narrow production signatures of bulk-generation? ISRC concentration asks: how many distinct legal entities route the copyright claims? Release cadence asks: are tracks uploaded in clusters consistent with automation? Catalog density asks: how concentrated is the set of companies credited as producers? Genre concentration asks: how narrow is the genre profile? The radar combines neighborhood structure with density. Cross-platform discrepancy asks: if this artist is real, where are they elsewhere online? Seven independent causal pathways from the same latent root.
No single measure can specify the ghost-artist class. But seven imperfect signals, when they all agree, constitute identification.
Each signal alone is unreliable. Audio variance cannot distinguish a ghost from a legitimate ambient specialist who has chosen a narrow palette and has stuck with it. ISRC concentration fails on indie artists using DistroKid, who all route through a single registrant by design. Release cadence could reflect algorithm-driven playlist-stacking strategies by legitimate artists. Catalog density fails on prolific classical composers whose rights are clearinghouse-managed through a single entity. Genre concentration fails on artists with a genuine stylistic focus. Each individual signal produces false positives. The neighborhood radar can rank an indie artist identically to a ghost. Cross-platform discrepancy flags emerging artists who simply have not yet broken through. None of these signals, taken alone, can specify the ghost-artist class. But the union of seven false-positive machines, when they all point to the same artists, is harder to dismiss as coincidence. That is where the identifying assumption—conditional independence given G—must hold.
The mechanisms are genuinely different. Each signal routes through its own causal pathway to G. The assumption is strong, largely unverifiable, and does fail on real data.
The identifying assumption is conditional independence: once you condition on whether an artist is actually a ghost, the noise in Signal 1 should be independent of the noise in Signal 4, independent of the noise in Signal 5, and so on. The mechanisms are genuinely different. Audio variance comes from how many distinct sound profiles appear in a catalog. ISRC HHI comes from how many distinct legal entities the copyright routes through. Release cadence comes from upload timing patterns. Playlist entropy comes from the diversity of playlist features. If the assumption holds, joint agreement on the same artists is not mere coincidence; it is what you would expect when several independent measurement instruments are all sampling from the same latent state. This assumption is strong and largely unverifiable. It can fail if two signals pick up the same artifact of data construction, or if the signals are actually competing measures of a different latent variable. The paper surfaces one such failure in the proxy-data collinearity diagnostic.
Think of it like triangulation in physics: each tool has noise and bias, but three independent tools pointing to the same location constitute evidence of something real.
The identifying move is not back-door adjustment or instrumental-variable reasoning. There is no treatment, no counterfactual to estimate. It is closer to latent-class identification through pattern convergence. Build several imperfect operationalizations of the ghost-artist concept. Where they agree, that agreement itself is the identification. Think of it like triangulation in physics: each measurement tool has noise and bias, but if three independent tools all point to the same location, you have found something real. The framework here says: if audio variance, ISRC concentration, release cadence, catalog density, genre concentration, neighborhood radar, and cross-platform discrepancy all rank the same set of artists as high-suspicion, the joint convergence is evidence that you have identified a real underlying class, not an artifact of any single measure. Convergence—the intersection of seven imperfect signals—is the finding. It is also the most the data can claim.
Ghost artists route catalogs through far fewer distinct legal entities than organic artists. The effect is strong and highly significant on ground-truth data.
On the confirmed-ghost ISRC data, Mann-Whitney testing shows that ghost HHI is significantly higher than organic HHI, U = 90.0, p = 0.0026, rank-biserial r = 1.000. This means the distribution of copyright-registrant concentration is shifted toward the high end for ghosts. The effect is not marginal. In plain language, ghost artists route their catalogs through far fewer distinct legal entities than organic artists do. An indie artist using DistroKid might have ten tracks routed through ten different registrants if they worked with different session musicians or labels. A ghost artist uploading 600 tracks in bulk uses one registrant, or two, or three. The signal works on the ground-truth data. On the 114K-track Kaggle proxy, the differences are measurable but smaller in effect size, reflecting the noisier operational measures available without direct ISRC access. The convergence of all seven signals strengthens each individual signal's claim.
Careful empirical work surfaces where the framework's central assumption breaks. When two signals measure the same data-construction artifact, convergence becomes correlation.
The paper surfaces one critical place where the conditional-independence assumption breaks. The S5 sign-flip diagnostic found that on the Kaggle proxy data, genre concentration and playlist entropy are collinear. When both are added to a logistic model, the incremental AUC contribution of genre concentration is 0.0000. The signals should be conditionally independent given G; instead they are picking up the same pattern through the same artifact of how the proxy dataset was constructed. This failure does not invalidate the framework; it identifies a place where the framework's central assumption visibly collapses on the available data. A careful empirical strategy surfaces these breaks. The collinearity suggests that proxy-data results on Signals 2 and 5 should be interpreted with caution and that ground-truth ISRC data is essential for robust convergence claims.
Ghost artists leave fingerprints across production, licensing, release timing, and cross-platform presence because they are produced by a different operational pipeline. This is inference of a hidden category from its observable descendants, not causal-effect estimation. The identifying assumption is conditional independence; when that assumption holds, convergence is identification.
The thesis that threads through this case is that when no single field or audit trail can identify a latent class, convergence of multiple independent signals—each operating through a different causal mechanism—is the identifying strategy. Ghost artists leave fingerprints across production, licensing, release timing, and cross-platform presence because they are produced by a different operational pipeline. No one fingerprint is diagnostic. Their joint presence is. This is not a causal-effect estimate; it is inference of a hidden category from its observable descendants. The identifying assumption is conditional independence: the noise in each signal is uncorrelated with the noise in the others, given the latent class. When that assumption holds, convergence is identification. When it breaks—as in the collinearity diagnostic—careful empirical work surfaces the break and bounds the claim. Convergence works because it forces multiple independent mechanisms to agree.
Causal Inference · Chapter 23 · Case: GhostTrack