Sockpuppet Detection in Hebrew Wikipedia: Network Analysis & SNA
A graph-based investigation of whether the interaction structure of Wikipedia discussions carries a detectable signature of sockpuppetry beyond the text itself.
Does sockpuppetry leave a structural footprint in discussion networks?
The project deliberately explores a different signal from text-based authorship analysis. Wikipedia talk-page discussions are converted into directed interaction graphs, with users represented as nodes and reply relationships represented as edges.
The central question is whether discussions involving confirmed sockpuppets or their puppet-masters exhibit network structure that differs systematically from ordinary discussions.
Compared network measures
- In-degree and out-degree
- PageRank
- Degree and eigenvector centrality
- Closeness and betweenness centrality
- Reciprocity and density
- Connection-graph entropy
From Wikipedia HTML to directed interaction graphs.
The system scrapes confirmed sockpuppet information, parses discussion-page HTML to recover reply relationships, constructs directed NetworkX graphs, and labels participants as ordinary users, confirmed sockpuppets, or confirmed puppet-masters. Duplicate discussions and historical parsing inconsistencies are filtered before analysis.
The strongest signal came from puppet-masters, not from puppets alone.
Infected vs. uninfected
The direct test of puppet-containing discussions against clean discussions was largely negative. Only closeness centrality reached statistical significance (p ≈ 0.018), with a very small effect size (Cohen’s d ≈ −0.19).
Masters vs. uninfected
Every examined whole-graph metric reached statistical significance in the master-containing comparison, but the effect sizes remained small. The largest reported separations were around |d| ≈ 0.27–0.31.
Structural interpretation
Master-containing discussions tended to be sparser, less reciprocal, and less interconnected. The infected group generally lay between the master and uninfected groups.
Important limitation
Several graph metrics were mathematically redundant, and averaging node-level values across an entire discussion can dilute precisely the local signal the study is trying to detect.
Moving from graph averages to individual positions revealed substantially larger effects.
The project then aggregated metric values by individual user across discussions. Among 653 users represented in more than five discussions, confirmed sockpuppets were not significantly different from ordinary users on the eleven tested user-level metrics. Puppet-masters, however, separated much more clearly.
PageRank
Masters vs. regular users: p ≈ 0.0004, Cohen’s d ≈ 0.73, the largest reported individual-level effect.
Eigenvector centrality
Masters vs. regular users: p ≈ 0.003, d ≈ 0.71.
Degree & closeness centrality
Degree centrality reached d ≈ 0.62 (p ≈ 0.002), while closeness centrality reached d ≈ 0.64 (p ≈ 0.009).
Sockpuppets vs. masters
Connection-graph entropy was the principal significant separator reported between the two groups (p ≈ 0.042, d ≈ −0.63).
A structural tendency, not a standalone detector.
The project’s whole-graph results do not support a claim that a discussion can be reliably classified as sockpuppet-infected from averaged network metrics alone. The more informative result is that confirmed puppet-masters occupy systematically different network positions at the individual level.
Several explanations remain plausible, including selection effects in how masters are identified, different functional roles for masters and puppets, participation-frequency differences, and signal dilution through averaging. The report explicitly treats these explanations as hypotheses rather than established mechanisms.
Next research directions
- Node-level features
- Distributional rather than averaged metrics
- Subgraph-level signatures
- Participation-frequency controls
- Targeted role analysis
- Combined structural and textual evidence
Final report.
The student final report for this project, as a PDF.
Network structure complements rather than duplicates NLP-based sockpuppet research.
The project establishes a separate structural research direction: instead of asking whether two accounts write alike, it asks whether suspicious identities occupy different positions in the interaction network.