NETWORK PRIVACY / FORENSICS / TRAFFIC ANALYSIS

P2P Anonymity: Digital Investigations and End-to-End Traffic Correlation

Moving from “not leaking your residential IP” to “making activities harder to link.”
A reorganized research report, from an overview and underlying principles to defensive architectures and verification methods.

Not seeing the real IP does not mean correlations cannot be established.
Source version: 2026-08Consulted / reviewed: 2026-10-03Main text + research notes + original-text cross-reference
2Core security boundaries
7Anonymity properties
12Reorganized topics
63Original sections retained
Article contents

Reorganized original textPrimary-source verificationReview notes / research proposals

How to read: start with Chapters 01–02 for an overview, then continue for the technical details. The main text combines repeated arguments while retaining the original terminology, research questions, and a cross-reference for all 63 sections. New derivations, limitations, and recommendations are labeled separately; the full source text can be expanded in Appendix C.

This document discusses lawful data exchange and authorized privacy research. All diagrams are conceptual redrawings; no network tests were run and no product anonymity scores were produced for this edition.

01

OVERVIEW · Two distinct security boundaries

Start with the big picture: which link do you want to break?

The real question in this report goes beyond “does turning on a VPN help?” It asks: What can others observe, which events can they link, and can they ultimately attribute the activity to a person? The original's 63 sections cover protocols, system isolation, digital investigations, and anonymous-communication research. This edition first places them on a shared map of the problem.

LAYER A / Leak prevention

Does the residential IP go out directly?

This concerns which interface P2P packets use, and whether they fall back to the physical network when the VPN disconnects.

P2P → VPN; VPN disconnects → external traffic stops

Tools: interface binding, firewalls, Network Namespace.

LAYER B / Correlation resistance

Can different activities be recognized as belonging to the same person?

Even after the exit IP is replaced, timing, traffic shape, accounts, and device evidence may still create links.

Activity + timing + metadata → same operator?

Research: traffic shaping, mixing, cover traffic, and cross-layer unlinkability.

This edition's central assessment: Layer A is an engineering boundary that can be clearly defined and tested; Layer B is a security problem dependent on observation capabilities, data distributions, and attack models. Passing tests for the former does not automatically guarantee the latter.

Original text §9, §25, §58, §63
Your deviceApps, accounts, files
Tunnel entryResidential IP / encrypted traffic
VPN or anonymity networkRelays and trust boundaries
Exit and destinationExternal activity / protocol data
System map: investigative information can come from any layer; it need not come from breaking encryption.

Three reading questions

What do you want to answer? Where should you start? What you should take away
Will P2P leak the residential IP? Chapters 03 and 05 Network boundaries and failure acceptance criteria
What can VPN, Tor, I2P, and DAITA each do? Chapters 02, 06, and 07 A capability comparison with explicit assumptions, rather than an anonymity ranking
How should end-to-end traffic-correlation defenses be researched? Chapters 08–11 Testable architectures, attack baselines, and cost metrics
Scroll the table horizontally →
02

THREAT MODEL · First clarify who is protected and against whom

Anonymity is not a switch: seven goals and observers

The original breaks anonymity into seven Security Properties. Retaining this classification is more meaningful than assigning a single score such as “90% anonymous.”

Security property Plain-language question Distinctions to preserve
Real-IP Confidentiality/ Real IP confidentiality Can trackers and peers see the home or company WAN IP? Replacing the exit with a VPN does not make activities unlinkable.
Identity Anonymity/ Identity anonymity After observing an activity, is the operator's identity known? An address, an account, a device, and a natural person are different things.
Activity Unlinkability/ Activity unlinkability Can activities A, B, and C be attributed to the same person? Long-term activity clusters may be formed even without a real name.
Cross-Layer Unlinkability/ Cross-layer unlinkability Can website accounts, browsers, and P2P activities be linked? A VPN does not remove application-layer accounts or device records.
Fingerprint Minimization/ Fingerprint reduction Can software versions, Peer IDs, or protocol extensions identify the client? Reducing fields does not hide the network endpoint from the other party.
Traffic-Analysis Resistance/ Traffic-analysis resistance Can the two ends be matched using packet timing, direction, and size? Statistical signals may exist without decrypting the content first.
Longitudinal Anonymity/ Long-term anonymity After hundreds of activities accumulate, does the candidate set keep shrinking? Safety in one instance does not imply long-term safety.
Scroll the table horizontally →
Original text §4, §8

Observer capabilities matter more than product names

The following analysis matrix is organized from the original. It is not a fixed list of privileges: actual visibility depends on routing, encryption scope, software behavior, and whether other data has been obtained.

Observer Under a typical “properly isolated VPN + BitTorrent” model What this vantage point alone usually cannot establish directly
Peers / trackers in the swarm VPN exit endpoint, participating torrent, protocol and timing information Which residential user is behind the exit
Local network or access ISP The user's connection to the VPN, timing, packet sizes, and total volume Reading the complete torrent content from the tunnel contents alone
VPN server / its controller Entry, exit, or live forwarding relationships, depending on the architecture “Visible” does not mean “permanently recorded in history”
An adversary able to observe both ends Traffic sequences at entry and exit, which can be used for attempted matching A matching score is not proof of a natural person's identity
Investigators with endpoint or account data Device files, application state, logins, account relationships, and similar data Data sources and attribution must be checked individually; their existence cannot be assumed
Scroll the table horizontally →
Original text §5, §6, §26, §60[R01][R02][R09]
H(U | O)

U is the candidate user, and O is the observations obtained. The original describes anonymity as “how much uncertainty about identity remains after observation,” rather than treating encryption-algorithm strength as anonymity itself.

03

PROTOCOL · Find peers before exchanging content

Where does BitTorrent leave observable data?

BitTorrent was designed to efficiently find other nodes and exchange data, not to hide participants. First distinguish indexes / indexing websites, trackers / coordination services, and peers that actually exchange content: they are not the same service, and need not be controlled by the same entity.

Original text §1, §2
DISCOVERY

Finding others

Trackers, DHT, and PEX provide or exchange endpoint information; LPD handles local peer discovery.

TRANSFER

Connecting and exchanging pieces

Actual peer connections reveal external endpoints and exchange behavior; routing determines whether the peer sees the residential IP or an exit IP.

HISTORY

Repeated observations build a history

Endpoints, Infohashes, timestamps, and client information can be assembled into activity records, but the reliability of each record still requires assessment.

Three key discovery mechanisms

Mechanism What does the protocol do? Visibility and limitations
UDP Tracker/BEP-15 An announce can carry info_hash, peer_id, downloaded, left, uploaded, event、port。 The tracker can also observe the packet's source endpoint. Fields are client-reported and should not all be treated as verified facts.
DHT/BEP-5 get_peers(info_hash) queries for peers; announce_peer announces participation. Token validation is tied to the querying source IP; nodes accepting an announce store the source IP and the announced / inferred port.
PEX/BEP-11 Connected peers exchange contact information for other peers. It adds discovery paths without removing the visibility of existing connection endpoints.
Scroll the table horizontally →
[R01][R02][R03]
(IP, port, Infohash, timestamp, observation type)

This edition recommends this representation of an “observation record.” Port and observation type are included to avoid conflating tracker listings, announces, and actual piece exchanges as the same kind of evidence.

What does disabling DHT / PEX actually change?

It reduces specific peer discovery surface, but as long as trackers or direct peer connections are still used, the other side can still see the connection's network endpoint. Keeping the residential IP out of the swarm requires a complete network boundary, not merely disabling discovery features.

What does Anonymous Mode actually change?

qBittorrent officially describes it as reducing identifying-information exposure. Actual behavior varies with qBittorrent and libtorrent versions and torrent type. An older version's behavior of “disabling DHT” cannot be applied to every version, and Anonymous Mode is not a substitute for a VPN or a fail-closed firewall.

[R04]Original text §2, §4
04

INVESTIGATION · Correlation is neither conviction nor decryption

From network clues to identity attribution: what is still missing?

The original describes modern investigations as building a timeline from Network Evidence, Service Evidence, and Endpoint Evidence, followed by Entity Resolution, candidate ranking, and corroboration. This is an analytical framework, not a claim that every case can obtain all of the data.

Original text §5, §6, §7
Observe eventsNetwork / service / endpoint
Build a timelineAlign sources and uncertainties
Form candidate linksGraph models and hypotheses
Independent corroborationValidate or rule out alternative explanations
Missing links must remain “unknown”; a plausible story cannot fill them in as confirmed facts.

4.1 Graph models: nodes are not people, and edges are not proof

The original uses G = (V, E) to represent attribution. Nodes may be people, accounts, devices, IPs, VPN exits, sessions, torrents, payments, or timestamps; edges may be connected_from, announced, same_device, temporally_correlated and similar relationships.

Review note: Each edge should additionally record its source, time range, uncertainty, confidence, and whether it is inferred. Sharing an exit IP establishes only a candidate relationship; it does not justify merging entities into one person.

4.2 Bayesian updating: do not count the same evidence twice

P(H | E₁, E₂, …, Eₙ)

H is the hypothesis that a candidate performed the activity. Multiple observations may change confidence in it, but different E values may stem from the same observation or a common cause.

Three reports exported from the same tracker dataset are not three independent pieces of corroboration. The original specifically cautions against unconditionally assuming P(E₁,E₂ | H) = P(E₁ | H) × P(E₂ | H). This edition assigns no identity probabilities without an empirical basis.

4.3 Intersection attacks: long-term activity shrinks the candidate set

S* = S₁ ∩ S₂ ∩ … ∩ Sₙ

If each activity requires the operator to be online, long-term observation can progressively exclude people who fail that condition. However, errors in offline-status observations, missed captures, and alternative connections all affect the intersection; the algorithm does not necessarily converge to one person.

[R11]Original text §7, §8

4.4 Historical evidence and misattribution in P2P monitoring

In 2010, Spying the World from Your Laptop reported collecting approximately 148 million IPs and roughly 2 billion observations related to content copies over 103 days. This demonstrates that IP–content relationships could be built at scale at the time; IP counts cannot be treated as counts of distinct natural persons, nor extrapolated to any present-day investigator's complete coverage.

In 2008, Why My Printer Received a DMCA Takedown Notice demonstrated the risk of misattribution in monitoring methods. Thus “listed by a tracker,” “connection completed,” “verifiable content exchanged,” and the later question of “which person operated it” must be distinguished. These are different levels of claims; a single IP record does not establish all of them.

[R12][R13]Original text §3
CASE STUDY / Primary announcement and unknowns

The 2026 Nyaa / CODA / JHA case

CODA's announcement states that on July 28, 2026, Kyoto Prefectural Police arrested a person suspected of first distributing NHK content through Nyaa. It explains that CODA investigated Nyaa under METI-supported CBEP, obtained relevant information using an analysis tool developed by the Japan Hacker Association, and police then identified the suspect through their investigation.

The announcement supports the existence of the investigation and tool; it does not disclose technical details sufficient to reconstruct the algorithm. This edition does not present CDN records, cookies, VPN logs, a particular traffic-correlation algorithm, or “no swarm monitoring” as confirmed methods. The original's “index + tracker + timing + historical observations” is retained only as a possible analytical hypothesis.

[R14]Original text §18
05

ENGINEERING · Rule out bypasses before discussing stronger anonymity

The first testable boundary: fail-closed network isolation

The original's most actionable requirement is: All outbound P2P data flows must pass through the VPN; when the VPN is unavailable, they must not fall back to the ordinary WAN. This is a security invariant, not just a UI showing “connected.”

∀ p ∈ outbound P2P data flows: allowed_paths(p) ⊆ VPN tunnel

This refers to the application's data plane. The physical NIC still needs to transmit encrypted outer VPN packets; “zero traffic on the physical NIC” is not the correct acceptance criterion.

VPN unavailable ⇒ no outbound P2P traffic bypasses the tunnel

An interface still existing or an icon remaining lit does not establish tunnel availability; the actual target is direct-connection fallback after failure.

Original text §9, §10, §11, §12, §54

5.1 From weaker to stronger: more than checking extra options

Layer What does it provide? What still needs verification?
Routing-only VPN Routes the ordinary default route into the tunnel. Whether more-specific routes, exception routes, or system changes create bypasses.
Application Binding Binds qBittorrent to the designated VPN interface. Whether interface recreation, version-specific behavior, and additional outbound programs are covered.
Firewall Fail-Closed Explicitly blocks protected traffic that does not pass through the tunnel. IPv4 / IPv6, rule initialization and reloads, all exits, and permitted exceptions.
Network Namespace Gives P2P a network space without an ordinary WAN path. Whether management interfaces, veth, DNS, or host privileges reopen a bypass.
Scroll the table horizontally →
[R05][R06][R07][R08]
P2P NETWORK NAMESPACE
qBittorrent / P2P Engine
↓
wg0 · the only permitted Internet interface

Local loopback may exist, but an ordinary WAN fallback must not.

HOST / PHYSICAL NAMESPACE
WireGuard UDP socket
↓
eth0 / wlan0 → VPN server

WireGuard's official model sends outer UDP from the namespace where the interface was created.

Redrawn from WireGuard’s official Network Namespace architecture. This is a Linux networking model, not a feature present under the same name on every operating system.

5.2 Why is inspecting the routing table alone insufficient?

Bypassing Tunnels (USENIX Security 2023) demonstrated VPN traffic leaks caused by abused routing exceptions; TunnelVision (CVE-2024-3661) demonstrated bypass risks from DHCP Option 121 and more-specific routes. They support the claim that “routing state is not complete isolation,” not that every VPN client today still has the same vulnerabilities.

5.3 Include IPv6, DNS, and management channels in the boundary

IPv4 through the VPN with IPv6 going out directly is still a failure. DNS also needs an explicit path. A leaked DNS query may reveal a tracker domain, but does not mean the ISP directly obtains the torrent's Infohash. Outbound requests through a local DNS proxy, WebUI, or additional search / update components require separate path verification.

[R06][R07][R08]Original text §13, §14
06

ARCHITECTURES · Examine paths, then trust distribution

VPN, Seedbox, Tor, I2P: avoid a single ranking

Section 57 of the original compares many techniques using “high / medium / low,” while noting that this is not a unified benchmark. This edition retains the comparative purpose, replacing easily misread overall scores with “function, assumptions, and uncovered risks”.

Architecture / measure Primary function Required assumptions Guarantees that should not be claimed
BitTorrent protocol encryption Protects the readability of certain content or protocol exchanges. The actual protocol and configuration are correct. Does not hide peer endpoints; it is not an anonymity protocol.
qBittorrent Anonymous Mode Reduces some client-identifying information. Version and torrent type behave as expected. Cannot replace IP routing isolation.
VPN Participates externally using the VPN exit IP. All relevant traffic enters the tunnel. Does not automatically provide activity unlinkability.
VPN + binding + firewall / namespace Reduces accidental direct connections and fallback risks. The boundary is complete and failure tests pass. Is not a defense against end-to-end traffic correlation.
Seedbox Moves swarm participation to a remote host. The home side uses only the intended management / retrieval channels. Trust shifts to hosting; accounts, endpoints, and activities may still be linked.
Same-provider multihop Separates entry and exit observation points. Intermediate paths and administrative arrangements match the model. Does not distribute trust across multiple operators, nor necessarily mix traffic.
VPN+DAITA Changes tunnel-packet shapes and some traffic fingerprints. Supported application / path, with a matching attack model and evaluation. Is not proven resistance to global, long-term correlation.
Tor Distributes path information across multiple relays. Correct application usage and the applicable threat model. Does not guarantee resistance to an adversary observing both ends simultaneously.
I2P-native P2P Uses overlay Destinations and tunnels, rather than clearnet IPs as application peer identities. Both the application and peer discovery stay within the I2P model. I2P still acknowledges timing, intersection, and other limitations.
Loopix/Mixnet Combines mixing, delay, and cover traffic. The protocol, security model, and load conditions are followed. Results for messaging systems cannot be applied directly to high-speed BitTorrent.
Scroll the table horizontally →
Original text §15, §16, §17, §57[R04][R06][R09][R11][R17][R18][R23]

Seedbox: changes where the data resides

Home → HTTPS/SSH → Seedbox → Swarm

It separates the residential network from the swarm, without eliminating the remote host's data, accounts, or management relationships.

I2P: changes the application communication model

App → I2P Destination/Tunnels → I2P Peer

It is not an ordinary clearnet swarm with a VPN exit simply wrapped around it. Overlay identity and underlying routing visibility must also be distinguished.

This edition's conclusion: For residential-IP isolation in ordinary clearnet P2P, verify the VPN boundary first. For P2P within an anonymity network, study its native overlay. For unlinkability against strong observers, discuss mixing, cover traffic, and costs rather than merely comparing VPN hop counts.

Original text §58, §62, §63
07

DEPLOYMENT · Separate policy, infrastructure, and traffic defenses

Mullvad and DAITA: six mechanisms with six different meanings

The original positions Mullvad as a privacy-oriented VPN, not an “untraceable communication system.” That positioning should be retained. The policy and deployment information below comes from primary documents consulted on 2026-10-03; provider statements, paper experiments, and this edition's inferences are labeled separately.

Mechanism Role in the original model Limitations and verification notes
Shared exit Reduces direct correspondence between the residential IP and activity. A shared exit is not a complete traffic mixer.
Numbered account Does not use traditional email / username registration. Accounts, WireGuard configuration, and payment data must still be understood under their respective policies.
No activity logs Reduces the possibility of retrospective lookup using server-side history. Does not mean no live connection state, no payment data, or no aggregate monitoring.
RAM-only infrastructure Reduces the risk of persistent disk remnants. Does not mean “traffic in memory or on the network can never be observed.”
Multihop Places entry and exit at different locations. A single provider does not distribute trust among independent operators.
DAITA Perturbs packet sizes, background traffic, and traffic shapes. A defense against traffic analysis, not a replacement of WireGuard with stronger content encryption.
Scroll the table horizontally →
Original text §19, §20, §21, §22, §23, §24, §59[R15][R16][R17][R18]

7.1 No-log and payment data: where is the missing link?

Mullvad's policy states that it does not retain user traffic, DNS, connection timestamps, IP, or user bandwidth logs, and describes live state used to verify concurrent connections. The same page also lists account settings, payment, and aggregate system data. Some payment records may support “person → payment / account,” but cannot thereby fill the missing historical link from “account → exit activity at a particular time”.

This is both the value and limit of data minimization: retaining less server-side data reduces retrospective traceability, but does not prevent third-party live observation or erase endpoint data.

[R15]Original text §20, §21, §22

7.2 DAITA v2: fixed size does not mean fixed rate

Official DAITA materials describe fixed packet sizes, random background traffic, and pattern perturbation. The v2 announcement further describes dynamic defense configurations assigned per VPN connection and reduced use of dummy packets. “Approximately halved” refers to the kind of dummy packets discussed in the announcement, not a halving of all traffic overhead. The unit of dynamic configuration cannot simply be equated with each BitTorrent peer flow.

[R18][R19]

Constant packet size

Pads packets of different sizes to a fixed size, reducing size signals; packet timing and counts may still differ.

Constant transmission rate

Maintains a sending rhythm even without real data. This is a stronger traffic-shaping requirement with a different cost structure; it does not automatically follow from fixed sizes.

7.3 The 2026 paper: deployment evidence with conditional defenses

Ephemeral Network-Layer Fingerprinting Defenses studies using different defenses for each connection and reports integration with WireGuard and actual deployment at Mullvad. Its evaluation includes circuit, website, and video fingerprinting; this does not amount to completed end-to-end correlation testing for all P2P traffic.

Section 5.3 of the same paper notes that, in a particular closed-world simulation, several padding-only defenses offered limited remaining protection once attackers had enough training data and time. Defenses introducing blocking / delay still helped, but their effectiveness also declined. This is a result for a specific model: it must not be rewritten as “all padding is useless,” nor omitted while citing only successful deployment.

[R20]
08

TRAFFIC CORRELATION · From intuition to a mathematical model

End-to-end traffic correlation: comparing shapes without reading content

Suppose an observer obtains sequence X on the user-to-VPN side and several candidate sequences Y₁, Y₂, … at the exit. The question is: Which exit candidate might correspond to this entry activity? Encryption can make content difficult to read, but does not automatically remove packet timing, size, direction, bursts, idle periods, or duration.

Original text §25, §26[R09]
X = {(tᵢ, sᵢ, dᵢ)};Y = {(t′ⱼ, s′ⱼ, d′ⱼ)}

t is time, s is size, and d is direction. Define the observation point first: outer tunnel packets, inner-tunnel IP traffic, and a single application flow cannot be mixed within one formula.

Entry XBurst → gap → burst
Exit YDelayed and perturbed, but potentially still similar
An educational illustration, not captured traffic or a real attack success rate for any product.

8.1 Basic time-window matching

xₖ = bytes(X, tₖ, tₖ + Δt)
yₖ = bytes(Y, tₖ, tₖ + Δt)
C(τ) = Corr(xₖ, yₖ₊τ)

Divide traffic into time windows and allow a time shift τ. A high score can rank candidates; it is neither proof of identity nor a calibrated probability of “confidence that this is the same person.”

Original text §27

8.2 Why does low-latency forwarding tend to preserve signals?

After real data arrives, a low-latency proxy generally must forward it quickly, preserving causal and temporal structure. The original illustrates this with a cumulative-data approximation:

Bᵧ(t + δ) ≈ Bₓ(t)

δ denotes forwarding delay. This is an intuitive approximation after aligning data and observation layers, not a strict law that “arbitrary entry bytes equal some exit flow's bytes.”

8.3 Passive observation and active perturbation

Passive adversaries compare existing signals at both ends; active adversaries may influence delay, loss, or rate to create recognizable changes on the other side. Their capabilities differ: defeating only a passive baseline does not establish equal safety against active adversaries.

8.4 Traffic fingerprinting is not traffic matching

Website Fingerprinting(WF) asks “which website does this encrypted traffic resemble?”; Flow Correlation(FC) asks “are these two traffic traces part of the same communication?” They may share features, but their labels, candidate-set sizes, and observation points differ. DeTorrent evaluates them separately, reminding us that their metrics cannot simply be interchanged.

Original text §31[R21]

8.5 Mutual information: retain the direction, limit the interpretation

The original uses Y = T(X) + N and I(X;Y) → 0 to express reducing dependence between the ends. This helps explain the research motivation. A formal experiment, however, should first define the random variables X and Y and whether shared traffic demand or external load creates dependence.

Original direction: min I(X; Y)
Review addition: evaluate I(M; O) or H(U | O)

M can be defined as the true entry–exit mapping, and O as the adversary's observations. This supplementary objective does not replace the original; it explicitly models the secret to be hidden. Estimates still depend on the model and data and are not unconditional security proofs.

Original text §28, §29, §52, §63
09

TRADE-OFFS · Consider latency, bandwidth, and anonymity goals together

Defensive techniques: what changes, and at what cost?

The original's Anonymity Trilemma intuition is that hiding “who has data when” generally requires sending cover traffic even without data, or waiting for more data to arrive before mixing. The former costs bandwidth; the latter costs time.

Comprehensive Anonymity Trilemma (2020) incorporates user coordination into a broader formal model and still derives limitations on anonymity and cost. This is a theoretical result under defined anonymity, adversary capabilities, and protocol models, not a slogan valid for every anonymity requirement. It does not negate the research value of improving systems against limited adversaries within practical budgets.

[R22]Original text §32, §33
Strong anonymityFewer exploitable correlation signals
Low latencySend real data sooner
Low bandwidth overheadFewer cover / dummy bytes

9.1 Defense mechanism matrix

Technique Observed features changed Main costs / conditions Limitations that must remain explicit
Constant-rate shaping Flattens sending rhythms and some idle-period signals. Sends dummies while idle; queues when load exceeds the fixed rate. Starts, stops, overload, and policy changes may still leak; this is more than fixing packet size.
Adaptive padding Inserts cover traffic at specific times. Bandwidth, CPU, and potential congestion costs. Real-packet timing may remain insufficiently hidden; no deliberate delay does not mean no performance cost.
Ephemeral shaping Uses different defense configurations for different connections. Configuration distributions, compatibility, stability, and evaluation costs. Adversaries may learn the defense distribution; randomness alone is not a security proof.
Mixing + delay Delays and reorders messages from multiple sources to weaken timing mappings. Delay, buffering, and protocol design. Requires enough mixing partners and handling of active attacks.
Multi-user aggregation Real traffic from multiple people provides mutual cover. Enough concurrently online honest users and mixing schedules. Shared NAT or a shared exit alone is not a cryptographic mix.
Traffic splitting/Multipath Distributes observations across multiple paths. Multipath availability, ordering / reassembly, and additional coordination. Partial observers may find matching harder; global observers may reaggregate.
Egress transformation Attempts to change the topology or timing of exit flows. Cooperative endpoints, a reassembly layer, or a new application protocol. Ordinary TCP connections cannot be arbitrarily rewritten without handling their state semantics.
Scroll the table horizontally →
Original text §34, §35, §37, §38, §40, §41, §42[R20][R23][R24]

9.2 The intuitive cost of a fixed rate

Real rate < R₀: dummy traffic fills the remaining capacity
Real rate > R₀: without dropping data, queues and waiting times grow

This explains why a very low fixed rate cannot promise arbitrarily high throughput, low latency, and complete traffic hiding simultaneously.

Original text §34

9.3 DeTorrent: a learnable padding-only defense

DeTorrent (PoPETs 2024) uses competing neural networks to generate and evaluate cover-traffic strategies. Its abstract reports that, in its FC configuration at FPR = 10⁻⁵ , TPR fell to approximately 0.12, and actual traffic combined with Tor was also tested. This is a result for specific data, attackers, and settings, not a claim that each user has only a 12% risk. “Does not delay traffic” describes the padding-only strategy; it does not guarantee zero increase in end-to-end latency on real congested networks.

[R21]Original text §36

9.4 Loopix: mixing and cover traffic for a stronger model

Loopix (USENIX Security 2017) uses Poisson mixing, random delays, cover traffic, and loop messages, considering global network observers and active attacks. Overall message delays in the original paper were on the order of seconds. “Relatively low latency” is relative to mix systems, not equivalent to a VPN's interactive latency.

Review inference: This makes it more suitable for delay-tolerant messaging or asynchronous communication. Applying it to large, continuous P2P transfers requires separate evidence for throughput, scheduling, and cost; its security level cannot simply be transplanted.

[R23]Original text §38, §39
10

RESEARCH PROPOSAL · Retain the concept and add engineering boundaries

The original research proposal: a multi-flow mixing network with bounded delay

The goal is not to claim perfect anonymity, but to reduce exploitable entry–exit mapping signals within latency, bandwidth, and infrastructure budgets.

Multi-user ingressEntry Aggregator
Ephemeral shapingEphemeral Shaper
Bounded-delay mixingBounded-Delay Mixer
The original's first half: combine multiple users with dynamic traffic strategies.
Multiple relay pathsPath 1 / 2 / 3
Egress mixing / reassemblyEgress Mixer
DestinationDelivery according to the protocol
The original's second half: path distribution and egress transformation. Implementation must first establish which transformations preserve communication semantics.

10.1 Component responsibilities and open verification items

Component Responsibility in the original Requirements identified by this edition's review
Entry Aggregator Collects multiple users and flows. All traceable mappings cannot be centrally logged without inclusion in the trust model; define who can see sources and destinations.
Ephemeral Shaper Selects different padding, dummy, size, and burst strategies per flow / connection. Explicitly choose the unit of action; attacker training must cover new strategy instances, not just known parameters.
Bounded-Delay Mixer Reorders / waits within a maximum delay. Requires deadlines, maximum queues, and overload handling. With only one honest source, it cannot pretend a multi-user anonymity set still exists.
Multi-path Distributes traffic over different paths. Distinguish logical paths from genuinely independent observation points; compare partial observations with observations that can be reaggregated.
Real-Traffic Cover Prioritizes other people's real data as cover. Exclude the adversary's own traffic; node counts alone do not represent the effective anonymity set.
Egress Mixer/Transformation Merges, splits, reorders, or migrates traffic. Specify whether transformations occur inside the overlay or on external connections; handle state, reassembly, and endpoint compatibility.
Adaptive Privacy Controller Adjusts strategies among privacy, latency, and bandwidth. Explicit policies are needed for model failure, insufficient cover traffic, or resource exhaustion; silent degradation to direct connections is unacceptable.
Scroll the table horizontally →
Original text §43, §44, §45, §46, §47, §48, §49

10.2 The engineering issue to resolve first: Egress Transformation

The original's split / merge / shuffle / connection migration is a direction, not a completed transport-layer design. TCP is a byte stream with connection and ordering state. An ordinary peer TCP connection cannot be arbitrarily split into several external connections while assuming the remote peer remains unaware.

A minimum viable approach derived in this edition: Initially split, mix, and reassemble only within an owned overlay between ingress and egress proxies; at egress, restore the original external flow semantics. This first tests signal perturbation in the intermediate network, without claiming the exit topology is eliminated. Truly changing external connection structure requires compatible proxies, cooperative endpoints, or a new protocol.

[R24]

10.3 The privacy controller's objective function

minπ [ λₚ·Lprivacy + λₗ·Llatency + λᵦ·Lbandwidth ]
subject to:L ≤ Lmax,OB ≤ Bmax

The original constrained optimization is retained. Review note: normalize the terms first and explain the units and purpose of λ. Adding losses of arbitrary scales does not establish optimal privacy.

The original uses min Defender / max Attacker to express adversarial training. Implementation must first define the loss direction: if L_attack is the utility of attack success, the attacker maximizes and the defender minimizes it; if it is the attacker's classification-error loss, the direction is generally reversed. The min–max formula cannot stand alone without defining what L represents.

10.4 Three questions to answer before implementation

First: who must be honest? Can entry and exit collude? How much can a shared provider see? Does mixing require at least one honest relay or other honest users?

Second: what happens when the budget runs out? When the latency limit conflicts with the anonymity goal, should the system pause, reject new traffic, or reduce defenses with user consent? This should be a separate policy from “no direct connections when the VPN is unavailable.”

Third: what exactly does the prototype promise? Select the protocol, observation points, candidate-set size, and load type before claiming improvement over a baseline under a particular model. Do not promise unmeasured global anonymity.

Original text §45, §49, §50
11

VALIDATION · Measure leaks first, then matchability

How to verify: isolation tests and anonymity research are separate experiments

The original P2P Privacy Lab uses random test files, Linux ISOs, or self-generated data, measured on self-administered clients, VPNs, trackers, peers, and observers. This edition retains that scope: Collect traffic only in authorized environments; do not use real third-party users as test data.

Original text §53

11.1 Experiment A: Network Isolation Fault Injection

These tests support “no direct P2P egress under specified failures,” not “anonymity proven.” Observe the host's physical interface, VPN interface, and test tracker / peer simultaneously, verifying that only the expected exit endpoint participates.

Failure type Required scenarios Acceptance focus
Startup order P2P starts before the VPN, reboot, service restart. No direct announce or peer traffic before the boundary is ready.
VPN failure Disconnection, daemon crash, unresponsive server, reconnection. Must not switch to the ordinary WAN; an unresponsive tunnel does not necessarily remove its interface.
Route changes Default / more-specific route changes, interface recreation. No new fallback paths may appear.
Network switching Wi-Fi / Ethernet switching, suspend / resume. Rules, bindings, and namespaces remain valid.
Protocol bypasses Enable IPv6, switch DNS. Actual IPv4, IPv6, and DNS paths all meet requirements.
Management and extensions WebUI, search components, updates, helper processes. Additional check in this edition: every permitted exception has an explicit purpose and cannot act as an outbound proxy.
Scroll the table horizontally →
Original text §54[R05][R06]
Acceptance record template

The checkboxes below are for manual review only. They do not scan the device or indicate that tests have passed.

11.2 Experiment B: Traffic Correlation

Experimental group Purpose
Baseline: unmodified WireGuard Establish baseline matchability and performance.
A:Padding Measure the benefits / costs of adding dummies independently.
B:Ephemeral shaping Measure whether dynamic strategies resist retrained attackers.
C:Bounded delay Measure the effects of timing changes independently.
D:Multi-user mixing Compare low, medium, and high honest-user loads.
E: combined scheme Determine whether the combined effect truly exceeds individual measures relative to cost.
Scroll the table horizontally →

“DAITA-like” can only mean an experimental group inspired by its approach. Without the same implementation and configuration, the result cannot be called a Mullvad / DAITA benchmark.

Original text §55

11.3 Do not report accuracy alone

Metric How should it be interpreted? What must accompany it
TPR @ fixed FPR The proportion of true matches found. Threshold, number of negatives, attack model.
Precision/PPV Of candidates classified as matches, how many are genuine? Prior / base rate, candidate-set size.
ROC、PR curve Overall behavior across thresholds. The very-low-FPR region still needs enough samples.
Correlation / estimated mutual information May serve as auxiliary signals or research proxy metrics. Features, estimators, bias, and model limitations.
Latency P50 / P95 / P99 Typical and tail user experience. RTT, load, retransmissions, and baseline.
Goodput, completion time, resource usage The actual cost of completing work. Added in this edition; line bytes alone can be inflated by dummies.
Bandwidth overhead Additional bytes relative to real data. Same observation layer; whether headers, retransmissions, and idle cover are included.
Anonymity sets and long-term tests Whether candidates keep shrinking after multiple observed events. Distribution of honest users, not merely the total node count.
Scroll the table horizontally →
Original text §30, §51, §52, §56
OB = Bytesdefended / Bytesreal − 1

Use identical measurement intervals and an explicit definition of bytes. When real traffic is zero, this ratio is undefined; report idle-cover bytes per second separately.

11.4 Interactive calculator: even a low false-positive rate can yield many wrong candidates

Assume exactly one true match among N candidates. Then the expected number of true positives is TPR, and the expected number of false positives is (N − 1) × FPR. The following is this edition's mathematical example, not case data or measurements of an anonymity product.

Expected false positives10.00
Expected true positives0.90
PPV / positive predictive value8.26%

PPV = TPR ÷ [TPR + (N − 1) × FPR]. These are model values given one true match and a shared operating threshold, not a probability of a person's identity for an individual result. If the true match is outside the candidate pool, a separate open-world decision is needed.

11.5 Research-quality checks added in this edition

Avoid data leakage: Group training, validation, and test sets by user, workload, session, time, or environment so adjacent segments of the same activity do not appear in both training and testing. Test unseen websites, file sizes, congestion, and retransmission conditions.

Give attackers a fair chance: At minimum, compare simple statistical attacks with retrained learning-based attacks, and report known / unknown defense, passive / active, and short-term / long-term settings separately. The sufficiently trained results in the 2026 Ephemeral paper explain why this step cannot be skipped.

Low FPR requires enough negative examples: If zero false positives are observed among m approximately independent negative examples, the one-sided 95% upper bound under a binomial model is 1 − 0.05^(1/m), approximately 3/m. For example, when m = 300,000, the upper bound is still approximately 10⁻⁵. This is a statistical derivation added in this edition. Many pairs sharing the same flow are not independent; all pairs cannot be treated as equal amounts of evidence.

[R20]
12

REVIEW SUMMARY · Turn the discussion into auditable claims

Final assessment: what to retain, revise, and verify next

12.1 Retain the original's core conclusions

First, a VPN's direct value is network-identity separation. Combined with correct interface binding, firewalls, and namespaces, it can establish a strong, testable leak-prevention boundary, but does not guarantee cross-layer unlinkability.

Second, investigations need not break encryption. Metadata, timing, accounts, and endpoints can provide clues from multiple sources; every relationship must still distinguish observations, inferences, and independent corroboration.

Third, low-latency forwarding leaves correlation signals that can be studied. Strong defenses should directly measure adversary capabilities and latency / bandwidth costs, rather than merely adding VPN hops or encryption layers.

Fourth, the proposed composite architecture is a research starting point, not a product-capability claim. Ephemeral shaping, bounded delay, multi-user aggregation, real-traffic cover, and egress transformation all have explicit conditions awaiting testing.

Original text §58, §59, §60, §61, §62, §63

12.2 Key review notes in this edition

Original claim / likely interpretation Treatment in this edition Basis and status
“Anonymity” can be summarized as one capability. Retains seven properties and separates leak prevention from correlation resistance. Reorganization of the original classification.
Section 57's high / medium / low resembles a unified score. Rewritten as functions, assumptions, costs, and non-guarantees; the original table remains in the appendix. The original already noted it was not a unified benchmark; this edition avoids ranking misinterpretations.
No-log / RAM-only means no usable data exists. Distinguishes historical activity, live state, payment, configuration, and aggregate data. Official policy / infrastructure statements.
DAITA is deployed, so it comprehensively prevents end-to-end correlation. Distinguishes fingerprinting, flow correlation, and specific evaluations; adds limitations under sufficient training. Official feature page + 2026 paper.
Padding-only does not deliberately delay traffic, so it has no latency cost. Retains the algorithm-level description and adds congestion, bandwidth competition, and queueing effects. Paper model and this edition's engineering review.
Bᵧ(t+δ) ≈ Bₓ(t) means arbitrary packets are conserved across layers. Labeled as a conditional causal intuition; observation layers must first be aligned. Supplementary note on the original model.
I(X;Y) can prove anonymity on its own. Retains the research direction and adds definitions of secret M, observations O, and identity uncertainty. Modeling recommendation in this edition, not a measurement result.
Split / merge alone can change any exit flow. Separates internal overlay transformations from ordinary TCP's external semantics. TCP specification + engineering derivation in this edition.
The Nyaa case publicly disclosed a specific tracking algorithm. Retains only officially confirmed facts, labeling the rest as unknown or hypothetical. CODA primary announcement.
The Anonymity Trilemma rules out any possible improvement. States model conditions and preserves the research question of better trade-offs within practical budgets. 2020 paper + interpretation in this edition.
Scroll the table horizontally →
[R14][R15][R16][R20][R22][R24]
BOTTOM LINE

First remove packet bypasses,
then make activities harder to correlate.

Validate the first with system boundaries and fault injection. Evaluate the second with threat models, reproducible attacks, and cost curves. They should complement each other, rather than being conflated under one “anonymous” label.

12.3 Suggested research-delivery sequence

Stage Deliverables Conditions for proceeding
A / Definition Specify protected subjects, observers, paths, and cost limits. No undefined “complete anonymity.”
B / Isolation Topology, least privilege, failure matrix, and packet-observation records. No direct P2P egress observed under the listed failures.
C / Baseline Fixed datasets, observation points, statistical and learning-based attacks. Metrics and low-FPR sample sizes are explainable.
D / Ablation Add padding, delay, and mixing individually. Benefits and costs can be separated; unseen environments have also been tested.
E / Combination Multi-flow mixing prototype and honest / malicious user models. Protocol semantics are preserved, and security claims have explicit scope.
Scroll the table horizontally →
A

PRIMARY SOURCES

Sources and verification scope

The source is the user-supplied “Research Report on P2P Anonymity, Digital Investigations, and End-to-End Traffic Correlation” (2026-08, 63 sections). The main text's “Original §” links expand the corresponding text; [Rxx] links to primary sources in this section. Consulted: 2026-10-03.

Official policies are provider statements; paper conclusions depend on their experiments and threat models. This edition's derivations are not presented as external research results. The list does not claim to cover all recent research, and no independent product audits were conducted.

R01
Protocol specification

BitTorrent BEP-15 · UDP Tracker

Announce fields and tracker exchange format.

R02
Protocol specification

BitTorrent BEP-5 · DHT

get_peers, announce_peer, and source IP / token behavior.

R03
Protocol specification

BitTorrent BEP-11 · Peer Exchange

Endpoint-information exchange among connected peers.

R04
Official wiki

qBittorrent · Anonymous Mode

Feature purpose, version differences, and warning that strong privacy is not guaranteed.

R05
Official wiki

qBittorrent · How to bind your VPN to prevent IP leaks

Interface binding as an additional leak-prevention layer.

R06
Official architecture document

WireGuard · Routing & Network Namespaces

Separation between the namespace containing wg0 and the outer UDP socket.

R07
USENIX Security 2023

Bypassing Tunnels: Leaking VPN Client Traffic by Abusing Routing Tables

Research on routing exceptions and VPN bypasses; does not imply all current versions remain affected.

R08
Researchers / Leviathan

TunnelVision · CVE-2024-3661

DHCP Option 121 routing bypasses and mitigation discussion.

R09
Official support document

Tor · Limitations and remaining attacks

Traffic correlation from simultaneous entry and exit observation is outside Tor's defensive guarantees.

R10
Official support document

Tor · Can I use Tor with Torrent?

Torrent clients over Tor are not recommended.

R11
Official threat model

I2P · Threat Model

Threats including timing, intersection, Sybil, and traffic analysis.

R12
USENIX LEET 2010

Spying the World from Your Laptop

103 days, 148 million IPs, and 2 billion copies; historical observations, not current coverage.

R13
USENIX HotSec 2008

Why My Printer Received a DMCA Takedown Notice

Monitoring methods and misattribution problems.

R14
CODA · 2026-07-28

First Uploader Using “Nyaa” Arrested

Official confirmation of tools and investigative cooperation, without a reconstructible algorithm.

R15
Official policy · page updated 2026-06-17

Mullvad · No-logging of user activity policy

Activity logs, account settings, payment, live state, and aggregate data require separate interpretation.

R16
Official infrastructure announcement · 2023

Mullvad · Migration to RAM-only VPN infrastructure

Provider statement on migration to RAM-only.

R17
Official feature documentation

Mullvad · Multihop with WireGuard

Design purpose of separating entry and exit.

R18
Official feature introduction

Mullvad · DAITA: Defense Against AI-guided Traffic Analysis

Fixed packet sizes, background traffic, and pattern perturbation.

R19
Official announcement · 2025-03-28

Mullvad · DAITA version 2

Dynamic configurations and fewer dummy packets; not a uniform halving of all overhead.

R20
PoPETs 2026(1), pp. 426–449

Pulls et al. · Ephemeral Network-Layer Fingerprinting Defenses

Page 1: scope and deployment; §5.3 / printed page 434: padding-only limitations with sufficient training.

R21
PoPETs 2024(1), pp. 98–115

Holland et al. · DeTorrent

Adversarial padding-only defense, evaluating WF and FC separately; figures apply only to the paper's settings.

R22
PoPETs 2020(3), pp. 356–383

Das et al. · Comprehensive Anonymity Trilemma

A formal model incorporating user coordination into anonymity and cost limitations.

R23
USENIX Security 2017

Piotrowska et al. · The Loopix Anonymity System

Poisson mixing, cover / loop traffic, threat model, and second-scale message latency.

R24
IETF Internet Standard

RFC 9293 · Transmission Control Protocol

TCP connections, state, and reliable ordered-byte-stream semantics.

B

GLOSSARY

Quick glossary

Organized using the attachment's terminology; main-text sections govern detailed definitions and limitations.

Swarm
The set of nodes participating in the same torrent.
Peer
A peer node actually participating in data exchange or protocol communication.
Tracker
A service helping clients obtain other peers' contact information.
Infohash
Hash information identifying a torrent, not a user identity identifier.
DHT / PEX / LPD
Distributed hash table / peer exchange / local peer discovery: three different discovery paths.
Ingress / Egress
The entry / exit of a system boundary; the observation point must be specified.
Endpoint
May mean a network IP:port or a user's device; this article distinguishes them by context.
Metadata
Surrounding data that need not contain the payload, such as timing, endpoints, lengths, and protocol fields.
Attribution / Entity Resolution
Identity attribution / entity resolution: the analytical process of linking different observations to the same entity.
Fail-closed
Stop protected operations when the protection mechanism fails, without falling back to an unsafe channel.
Network Namespace
An isolated space for Linux network resources, not a complete host-security sandbox.
Padding / Dummy / Cover traffic
Padding / dummy traffic / cover traffic: related but not identical concepts; other real traffic may also provide cover.
Mixing / Aggregation
Mixing / aggregation: the former aims to break input–output mappings; merely putting traffic together in the latter does not guarantee the same effect.
WF / FC
Website fingerprinting / traffic correlation: distinct tasks of classifying content type and matching two ends.
TPR / FPR / PPV
Detection rate for true matches / false-positive rate for nonmatches / positive predictive value after a match decision.
H(U | O) / I(X;Y)
Conditional entropy / mutual information: mathematical measures of uncertainty and dependence; random variables and distributions must be defined first.
QoS / Goodput
Quality-of-service constraints / useful-data throughput; the latter does not count dummies as useful work.
C

SOURCE ARCHIVE / 63 SECTIONS

Original-section cross-reference and full text

Each original section is preserved in Markdown, including the original formulas, diagrams, and comparison tables. Review notes are not written back into the source. The collapsible content below is an archive of the original only, not a renewed endorsement of every claim by this edition.

PrefaceOriginal title and abstractOriginal L1–L62
# Research Report on P2P Anonymity, Digital Investigations, and End-to-End Traffic Correlation

**Version: 2026-08**
**Topics: BitTorrent / VPN / Mullvad / Tor / I2P / Traffic Correlation / Traffic Analysis Defense**

---

## Abstract

The core problem of modern online anonymity can no longer be reduced to “whether the real IP is hidden.” A fuller model must consider all of the following:

\[
\text{Network Identity}
+
\text{Protocol Metadata}
+
\text{Session Linkability}
+
\text{Temporal Correlation}
+
\text{Behavioral Fingerprint}
+
\text{Cross-Layer Identity}
+
\text{Endpoint Evidence}
\]

The goal of an anonymity system can be formalized as maximizing:

\[
H(U\mid O)
\]

Here \(U\) is the real user and \(O\) is all observations available to the attacker: maximizing the uncertainty about identity remaining after observing network, timing, account, device, protocol, and other information.

Investigators, conversely, seek to use:

\[
O_1+O_2+\cdots+O_n
\rightarrow
\text{Entity Resolution}
\rightarrow
\text{Attribution}
\]

to make:

\[
H(U\mid O_1,\ldots,O_n)
\downarrow
\]

Thus modern investigations usually do not directly “break a VPN” or “break WireGuard”; they gradually narrow the candidate set through **metadata correlation, timeline analysis, graph correlation, traffic analysis, account records, endpoint forensics, and multi-source evidence fusion**.

The core conclusion of this report is:

> **VPNs excel at breaking the direct link “residential IP → Internet activity,” but do not automatically provide activity unlinkability, traffic-analysis resistance, or long-term anonymity.**

Against a powerful attacker who can observe both ends of low-latency communication, there is still no practical general solution combining **low latency, low bandwidth overhead, and strong anonymity**. Tor's official documentation still explicitly acknowledges this limitation; the Anonymity Trilemma in anonymous-communication research further formalizes the underlying trade-off.

---

01The basic anonymity problem in P2POriginal L63–L116
# 1. The basic anonymity problem in P2P

## 1.1 BitTorrent is inherently not an anonymity protocol

BitTorrent's primary goal is:

> Efficiently finding other peers and exchanging data.

Not:

> Hiding who is participating in a swarm.

BitTorrent therefore natively has several observable surfaces:

```text
                  ┌── Tracker
                  │
BitTorrent Client ├── DHT
                  │
                  ├── PEX
                  │
                  ├── Peers
                  │
                  └── Local Peer Discovery
```

A standard UDP tracker announce can include:

- `info_hash`
- `peer_id`
- `downloaded`
- `left`
- `uploaded`
- `event`
- listening port

The tracker can also naturally see the source network endpoint.

A tracker can therefore readily form:

\[
(IP,\ InfoHash,\ Timestamp)
\]

Or even:

\[
(IP,\ InfoHash,\ Uploaded,\ Left,\ Event,\ Timestamp)
\]

This is normal BitTorrent protocol design, not a vulnerability.

---

02Exposure surfaces of DHT and PEXOriginal L117–L154
# 2. Exposure surfaces of DHT and PEX

BitTorrent DHT (BEP-5) is a distributed peer-discovery system.

One of its core operations:

\[
get\_peers(info\_hash)
\]

obtains peer contact information for a torrent; `announce_peer` announces participation in that torrent to the DHT.

BEP-5 also requires the token used by `announce_peer` to be bound to the querying source IP. The queried node ultimately stores that source IP and port under the corresponding infohash.

DHT can therefore be understood as:

> **A distributed `infohash → peer endpoint` discovery database.**

PEX exchanges other peers' information among peers with established BitTorrent connections, forming another peer-discovery surface.

Thus:

```text
DHT OFF
PEX OFF
```

can reduce the public discovery surface, but does not make it anonymous:

```text
Tracker → can still see the endpoint
Peers   → can still see the endpoint
```

The primary determinant of “whether the residential IP is exposed” remains the network routing boundary.

---

03Large-scale P2P monitoring is no longer theoreticalOriginal L155–L189
# 3. Large-scale P2P monitoring is no longer theoretical

**Spying the World from Your Laptop**, presented at USENIX LEET 2010, demonstrated that a single machine and BitTorrent's public infrastructure could build large-scale, long-term mappings between users and torrent activity.

Over 103 days, the study collected approximately:

- 148 million IPs
- Roughly 2 billion content-copy observations

It also reported inferring the initial content providers for many newly observed torrents.

An important study also points in the opposite direction.

At HotSec 2008:

**Why My Printer Received a DMCA Takedown Notice**

showed that relying solely on IP information returned by trackers can cause false attribution, even leading to infringement notices for devices that never actually exchanged data.

Forensic evidence must therefore distinguish:

\[
IP\ listed\ by\ tracker
\]

from:

\[
IP\ actually\ exchanged\ verified\ content
\]

These have different evidentiary strength.

---

04Anonymity should be separated into distinct Security PropertiesOriginal L190–L332
# 4. Anonymity should be separated into distinct Security Properties

“Anonymous” is not a single Boolean.

A more complete model can be written as:

\[
P=
(P_{IP},
P_{identity},
P_{activity},
P_{cross},
P_{fingerprint},
P_{timing},
P_{longitudinal})
\]

## 4.1 Real-IP Confidentiality

The aim is:

```text
Tracker
Peer
DHT
Website
```

cannot see the residential / company WAN IP.

This is the layer VPNs address best.

---

## 4.2 Identity Anonymity

Even if an activity is visible, the observer does not know:

```text
Activity X
    ↓
Who?
```

---

## 4.3 Activity Unlinkability

Even when there are:

```text
Activity A
Activity B
Activity C
Activity D
```

it should not be easy to determine:

\[
A=B=C=D=\text{same actor}
\]

This is stronger than IP hiding alone.

---

## 4.4 Cross-Layer Unlinkability

Prevent:

```text
Browser Identity
      │
      ├── Website account
      │
P2P Identity
      │
      └── Tracker/DHT
```

from being proven to belong to the same entity.

---

## 4.5 Fingerprint Minimization

Reduce:

```text
client version
peer ID pattern
User-Agent
protocol extensions
OS/application behavior
```

and other software fingerprints.

qBittorrent's Anonymous Mode mainly belongs to this category. Its official documentation explicitly says that it does not provide strong anonymity by itself, but reduces the identifying information exposed by the BitTorrent client.

---

## 4.6 Timing / Traffic-Analysis Resistance

The aim is for the entry:

\[
X(t)
\]

and the exit:

\[
Y(t)
\]

to be difficult to identify as the same communication.

This is one of the hardest problems for Tor, VPNs, and most low-latency anonymity systems.

---

## 4.7 Longitudinal Anonymity

Anonymity is not only about:

> Whether one activity leaked information.

It is also about:

> After hundreds or thousands of activities, do historical observations gradually reveal the identity?

The real subject of study is therefore:

\[
H(U\mid O^{(1)},O^{(2)},...,O^{(n)})
\]

not just a single session.

---

05The modern investigator's modelOriginal L333–L373
# 5. The modern investigator's model

Modern digital investigations can be abstracted as:

```text
                 Observable Event
                       │
        ┌──────────────┼──────────────┐
        ▼              ▼              ▼
     Network        Service        Endpoint
     Evidence       Evidence       Evidence
        │              │              │
        ├──────────────┼──────────────┤
                       ▼
                    Timeline
                       │
                       ▼
                Entity Resolution
                       │
                       ▼
                Candidate Ranking
                       │
                       ▼
                 Corroboration
                       │
                       ▼
                   Attribution
```

Investigators do not necessarily need one decisive piece of evidence.

More commonly:

\[
E_1+E_2+E_3+\cdots+E_n
\]

all support the same hypothesis.

---

06Graph-based Attribution ModelOriginal L374–L431
# 6. Graph-based Attribution Model

This problem is well suited to a model of:

\[
G=(V,E)
\]

Nodes may include:

```text
Person
Account
Device
IP
VPN Exit
Session
Browser
Torrent
Infohash
Tracker Event
Server
Payment
Email
Timestamp
```

Edges may include:

```text
connected_from
logged_in_from
observed_at
announced
created
controlled_by
paid_for
same_session
same_device
temporally_correlated
```

The essence of investigation is to start with many disconnected components:

```text
A        B        C
```

and gradually find credible cross-component edges:

```text
A ───── B ───── C
```

ultimately forming an entity cluster.

---

07Bayesian AttributionOriginal L432–L467
# 7. Bayesian Attribution

Assume:

\[
H_A=\text{Alice is the operator of Activity X}
\]

Given evidence:

\[
E_1,E_2,\ldots,E_n
\]

continually update:

\[
P(H_A\mid E_1,\ldots,E_n)
\]

This aptly describes how many weak signals gradually build strong attribution.

In practice, however, one cannot simply assume:

\[
P(E_1,E_2|H)
=
P(E_1|H)P(E_2|H)
\]

Many pieces of evidence share common causes or come from the same observation source.

Otherwise, **double counting evidence** is likely.

---

08Intersection AttackOriginal L468–L523
# 8. Intersection Attack

Another important anonymity model is the intersection of candidate sets.

Assume:

\[
S_0
\]

is the set of all possible users.

Online at the first event:

\[
S_1
\]

At the second:

\[
S_2
\]

Continuing:

\[
S^*
=
S_1\cap S_2\cap\dots\cap S_n
\]

may yield:

```text
10000
  ↓
3000
  ↓
600
  ↓
80
  ↓
12
  ↓
2
```

I2P's own threat model explicitly includes timing, intersection, predecessor, Sybil, and traffic analysis among anonymity attacks that must be considered.

Thus:

> **Long-term, repeated, regular anonymous activity is itself an accumulation of risk.**

---

09The first line of P2P defense: Network IsolationOriginal L524–L548
# 9. The first line of P2P defense: Network Isolation

If the security requirement is only:

> **BitTorrent's real residential IP must never reach a tracker, DHT, or swarm.**

The most reliable engineering security invariant should be:

\[
\forall p\in P2PTraffic:
OutgoingInterface(p)=VPN
\]

And:

\[
VPN\downarrow
\Rightarrow
P2PTraffic=0
\]

That is, **fail closed**.

---

10VPN + Application BindingOriginal L549–L567
# 10. VPN + Application Binding

qBittorrent can bind itself to a designated VPN network interface.

Official documentation explicitly describes this as an extra leak-prevention layer beyond the VPN kill switch: when the VPN interface does not exist, qBittorrent should not switch to an ordinary physical interface.

Typical:

```text
eth0 = physical WAN
wg0  = VPN

qBittorrent
     │
     └── bind wg0
```

---

11Firewall Fail-ClosedOriginal L568–L597
# 11. Firewall Fail-Closed

Going further:

\[
P2P\ process
\land
oif\neq wg0
\Rightarrow DROP
\]

In other words:

```text
               ┌── wg0  → ALLOW
qBittorrent ───┤
               └── eth0 → DROP
```

This can protect against:

- VPN daemon crash
- routing mistake
- reconnect
- interface changes

and similar failure scenarios.

---

12Network NamespaceOriginal L598–L636
# 12. Network Namespace

A stronger architecture is:

```text
Physical Namespace
│
└── eth0 / wlan0
        │
        └── encrypted WireGuard socket


P2P Namespace
│
├── qBittorrent
└── wg0
```

The P2P namespace has no:

```text
eth0
wlan0
```

Thus, even if the VPN route disappears, there is no ordinary WAN fallback path.

WireGuard's official documentation directly demonstrates this model: a container / network namespace can contain only `wg0`, so its only available Internet route is WireGuard.

Compared with relying solely on:

```text
default route → VPN
```

this makes a kernel-level security invariant easier to establish.

---

13Why routing-only VPNs are less than idealOriginal L637–L656
# 13. Why routing-only VPNs are less than ideal

**Bypassing Tunnels**, presented at USENIX Security 2023, showed that many VPN clients base their security boundaries on routing-table manipulation, and routing exceptions can become a traffic-leak attack surface. The researchers successfully demonstrated routing-based VPN bypasses across multiple platforms and VPN clients.

TunnelVision / CVE-2024-3661 further demonstrated injecting more-specific routes through DHCP Option 121 to bypass tunnels for some VPN traffic; the researchers also listed Linux network namespaces as a strong mitigation.

Thus:

```text
VPN connected = true
```

does not necessarily imply:

```text
∀ traffic → VPN
```

---

14IPv6, DNS, and other bypassesOriginal L657–L689
# 14. IPv6, DNS, and other bypasses

P2P isolation must address all of:

\[
IPv4
\]

from:

\[
IPv6
\]

Otherwise, this may occur:

```text
IPv4 → VPN
IPv6 → physical WAN
```

DNS should also be within the same network boundary.

A DNS leak does not usually mean the ISP automatically knows the torrent `infohash`, but may expose:

```text
tracker.example.org
```

and other service-level information.

---

15SeedboxOriginal L690–L739
# 15. Seedbox

A seedbox changes the network location:

```text
Home
 │ HTTPS/SSH
 ▼
Seedbox
 │
 ▼
BitTorrent Swarm
```

Thus:

\[
ResidentialIP\notin Swarm
\]

But:

```text
Swarm → Seedbox IP
```

still exist.

A seedbox therefore provides:

> **network-location separation**

Not:

> perfect anonymity。

The trust boundary merely moves from:

```text
VPN provider
```

to:

```text
hosting / seedbox provider
```

---

16Tor and P2POriginal L740–L759
# 16. Tor and P2P

Tor is a low-latency anonymous communication network. Its main security idea is to prevent one relay from knowing both:

```text
User
+
Destination
```

However, Tor's official documentation explicitly states:

> If an attacker can observe both ends of a communication channel, no practical low-latency system currently reliably prevents correlation using timing and volume.

Tor's primary strategy is therefore not to eliminate correlation, but to:

> Reduce the likelihood that an attacker simultaneously obtains observation positions at both ends.

---

17I2POriginal L760–L783
# 17. I2P

I2P's BitTorrent model differs from ordinary clearnet BitTorrent.

Its anonymous identity does not directly use clearnet:

```text
IP:port
```

as the application peer identity, instead using Destinations and tunnel infrastructure in the I2P overlay.

This therefore is:

> **protocol-level anonymous P2P overlay**

Not:

> clearnet BitTorrent + VPN。

Still, I2P itself explicitly acknowledges timing, intersection, traffic analysis, and related attacks as part of its threat model.

---

18The Nyaa / CODA / JHA caseOriginal L784–L825
# 18. The Nyaa / CODA / JHA case

In July 2026, Kyoto Prefectural Police in Japan arrested a man suspected of sharing NHK and other content as a Nyaa first uploader.

CODA officially confirmed:

- CODA investigated Nyaa under the METI-supported CBEP program;
- The Japan Hacker Association developed an analytical tool;
- CODA used the tool to obtain information related to the suspect;
- Kyoto Prefectural Police investigated further based on that information and identified the suspect.

However, the official sources **have not disclosed the tool's algorithm**.

Therefore:

```text
CDN logs
Cookies
Login history
VPN logs
```

and similar specific methods cannot be treated as confirmed facts.

More plausible technical hypotheses include:

```text
Index metadata
        +
Tracker metadata
        +
Temporal correlation
        +
Historical observations
        ↓
Candidate identification
```

But these remain inferences.

---

19Mullvad's anonymity modelOriginal L826–L850
# 19. Mullvad's anonymity model

Mullvad is a privacy-oriented VPN, not a complete anonymous communication network.

It primarily addresses the direct link:

\[
ResidentialIP
\rightarrow
InternetActivity
\]

.

Its main mechanisms include:

- shared VPN exit
- numbered account
- no activity logs
- Multihop
- DAITA
- RAM-only server infrastructure

---

20Mullvad No-LogOriginal L851–L874
# 20. Mullvad No-Log

As of 2026, Mullvad's official policy states that it does not permanently retain:

- traffic
- DNS requests
- connection timestamps
- disconnect timestamps
- session duration
- IP addresses
- user bandwidth

Live state needed for connection validation is handled in temporary memory rather than retained permanently.

Mullvad thus seeks to avoid creating a retrospective connection table of:

\[
(SourceIP,\ Account,\ ExitIP,\ Timestamp)
\]

.

---

21Numbered accounts and paymentsOriginal L875–L910
# 21. Numbered accounts and payments

Mullvad does not require a traditional email / password identity.

Payment itself may nevertheless create another attribution surface.

For example, some payment methods have records of:

```text
Payment
   ↓
Mullvad account
```

, and the official data policy explicitly describes retention for each payment method.

So:

\[
Person\rightarrow Account
\]

can sometimes be established.

Without connection logs, however:

\[
Account
\rightarrow
ExitIP@time
\]

still lacks a direct historical mapping.

---

22Mullvad RAM-only InfrastructureOriginal L911–L934
# 22. Mullvad RAM-only Infrastructure

In 2023, Mullvad announced completion of its VPN infrastructure's RAM-only migration. VPN servers no longer relied on ordinary persistent disks for operation; its infrastructure audit also examined this architecture.

Its security significance is reducing the risk of residual data from:

```text
server seizure
   ↓
historical disk artifacts
```

.

But RAM-only:

\[
\neq
\]

“no attacker can observe traffic in real time.”

---

23Mullvad MultihopOriginal L935–L969
# 23. Mullvad Multihop

Ordinary VPN:

```text
User
 │
 ▼
VPN
 │
 ▼
Destination
```

Multihop:

```text
User
 │
 ▼
Entry VPN
 │
 ▼
Exit VPN
 │
 ▼
Destination
```

Mullvad officially states that one purpose of multihop is to make incoming / outgoing traffic correlation harder by distributing observation positions across different servers, locations, and providers.

This is still not Tor-style trust distribution among multiple operators.

---

24DAITAOriginal L970–L1006
# 24. DAITA

DAITA:

**Defense Against AI-guided Traffic Analysis**

is primarily about changing the observable shape of encrypted traffic, rather than increasing cryptographic strength.

Mullvad officially lists these main techniques:

- random background / dummy traffic
- constant packet-size style padding
- traffic pattern distortion
- dynamic per-connection configurations

DAITA v2 further reduces dummy-traffic overhead and gives different connections different tunnel-traffic characteristics.

Its main targets are therefore:

\[
TrafficFingerprint
\]

and:

\[
TrafficCorrelationSignal
\]

Not:

\[
AES/WireGuardCryptography
\]

---

25End-to-End Traffic CorrelationOriginal L1007–L1040
# 25. End-to-End Traffic Correlation

This is the report's most important technical problem.

Assume an ingress observation:

\[
X=\{(t_i,s_i,d_i)\}_{i=1}^{n}
\]

Where:

- \(t_i\):timestamp
- \(s_i\):packet size
- \(d_i\):direction

At the exit:

\[
Y=\{(t'_j,s'_j,d'_j)\}_{j=1}^{m}
\]

The attacker must solve:

\[
Match(X,Y)?
\]

That is:

> Which exit flow corresponds to the same communication as ingress flow X?

---

26Encryption cannot hide all metadataOriginal L1041–L1071
# 26. Encryption cannot hide all metadata

Even if:

\[
C=Enc_K(M)
\]

the attacker may still see:

```text
packet timing
packet count
packet size
direction
burst structure
idle intervals
flow duration
total volume
```

Thus:

\[
Encryption\neq Traffic\ Unobservability
\]

This is also why Tor officially acknowledges the end-to-end correlation threat.

---

27The most basic correlation attackOriginal L1072–L1112
# 27. The most basic correlation attack

Divide time into windows:

\[
x_k=
bytes(X,t_k,t_k+\Delta t)
\]

\[
y_k=
bytes(Y,t_k,t_k+\Delta t)
\]

Then calculate lagged correlation:

\[
C(\tau)
=
Corr(x_k,y_{k+\tau})
\]

If:

```text
Candidate A = 0.17
Candidate B = 0.10
Candidate C = 0.91
Candidate D = 0.05
```

C is clearly worth further analysis.

An actual modern classifier can of course use more complex feature representations, but the core information sources remain:

\[
Timing+Volume+Direction
\]

---

28Data conservation is the core problemOriginal L1113–L1137
# 28. Data conservation is the core problem

Define cumulative bytes:

\[
B_X(t)=
\sum_{i:t_i\leq t}s_i
\]

A low-latency proxy must transfer ingress data to egress almost immediately.

Thus, generally:

\[
B_Y(t+\delta)\approx B_X(t)
\]

Here \(\delta\) is network delay.

Real data cannot appear at egress before it has arrived at ingress.

Low-latency forwarding therefore creates a causality constraint in itself.

---

29The mutual-information modelOriginal L1138–L1170
# 29. The mutual-information model

An ideal anonymity system aims for:

\[
I(X;Y)\rightarrow0
\]

In other words:

> Knowing ingress traffic provides almost no additional information for inferring egress traffic.

Ordinary VPN / Tor can, however, be abstracted as:

\[
Y=T(X)+N
\]

Where:

- \(T\):relay/network transformation
- \(N\):jitter、congestion、packetization noise

Usually:

\[
I(X;Y)>0
\]

and it may be quite high.

---

30False positives and base ratesOriginal L1171–L1216
# 30. False positives and base rates

Traffic-correlation papers must look beyond:

\[
TPR
\]

They must also consider:

\[
FPR
\]

Assume:

\[
FPR=10^{-3}
\]

With:

\[
10^7
\]

candidate comparisons, there may still be:

\[
10^4
\]

false positives。

Thus:

> targeted hypothesis confirmation

from:

> Internet-scale identification

are very different problems.

---

31Passive and active correlationOriginal L1217–L1245
# 31. Passive and active correlation

### Passive

Only observes:

```text
Ingress
Egress
```

then performs statistical matching.

### Active

The attacker can inject into the flow:

```text
delay
drop
rate perturbation
```

to create a distinctive temporal pattern, then observe whether a corresponding pattern appears on the other side.

Active attacks are generally stronger than purely passive observation.

---

32Anonymity TrilemmaOriginal L1246–L1281
# 32. Anonymity Trilemma

A fundamental trade-off in anonymous communication can be written as:

### Strong anonymity

\[
I(X;Y)\rightarrow0
\]

### Low latency

\[
L\rightarrow0
\]

### Low bandwidth overhead

\[
B_{dummy}\rightarrow0
\]

Research shows that anonymous communication systems cannot simultaneously achieve:

> strong anonymity + low latency overhead + low bandwidth overhead。

The Comprehensive Anonymity Trilemma at PoPETs 2020 extends the impossibility result to a broader class of anonymity systems that includes user coordination.

This means the problem is not simply:

> “A good enough padding algorithm has not yet been invented.”

There are more fundamental limitations.

---

33Why does the trilemma arise?Original L1282–L1333
# 33. Why does the trilemma arise?

Assume the user is currently idle.

If the system does not want an observer to determine:

```text
User is idle
```

it must continuously send:

\[
DummyTraffic>0
\]

Thus:

\[
BandwidthOverhead\uparrow
\]

Conversely, Alice suddenly generates substantial traffic.

To keep output from immediately reflecting the burst:

### Method A

Delay the data:

\[
Latency\uparrow
\]

### Method B

Send large amounts of cover traffic routinely:

\[
Bandwidth\uparrow
\]

So:

\[
StrongPrivacy
\]

Some resource must inevitably be consumed.

---

34Defense 1:Constant-Rate ShapingOriginal L1334–L1400
# 34. Defense 1:Constant-Rate Shaping

One of the strongest but most expensive methods:

\[
R(t)=R_0
\]

Whether or not real traffic exists:

```text
REAL
DUMMY
DUMMY
REAL
DUMMY
```

The observer always sees:

```text
████████████████████████
```

It can substantially hide:

- burst timing
- packet count
- volume shape
- idle periods

However:

When:

\[
R_{real}<R_0
\]

then:

\[
BandwidthWaste=R_0-R_{real}
\]

When:

\[
R_{real}>R_0
\]

the queue becomes:

\[
Q(t)\uparrow
\]

which then leads to:

\[
Latency\uparrow
\]

This is the most intuitive embodiment of the Anonymity Trilemma.

---

35Defense 2:Adaptive PaddingOriginal L1401–L1428
# 35. Defense 2:Adaptive Padding

A more practical approach is:

\[
X'=X+N
\]

Here \(N\) is selective dummy traffic.

Instead of a permanent constant rate, add cover only to traffic patterns carrying high fingerprint information.

Advantages:

- Almost no increase in latency
- Controllable bandwidth overhead
- Deployable on existing VPN / Tor systems

Disadvantages:

\[
I(X;X')>0
\]

generally still holds.

---

36DeTorrentOriginal L1429–L1458
# 36. DeTorrent

DeTorrent, published at PoPETs 2024, uses competing neural networks to automatically design padding-only defenses.

Defender:

\[
D_\theta(X)
\]

Attacker:

\[
A_\phi(D_\theta(X))
\]

This forms a kind of adversarial optimization:

\[
\min_\theta\max_\phi L
\]

Its flow-correlation evaluation shows a substantial reduction in attack TPR at a very low FPR operating point, without delaying real traffic.

Its value is that:

> Padding policies are no longer designed entirely by manual heuristics.

---

37Defense 3:Ephemeral DefensesOriginal L1459–L1488
# 37. Defense 3:Ephemeral Defenses

A fixed defense can itself become a fingerprint.

For example, if every user uses:

```text
padding distribution = θ
```

the attacker can learn:

\[
P(X'\mid\theta)
\]

**Ephemeral Network-Layer Fingerprinting Defenses**, at PoPETs 2026, proposes:

\[
\theta_i\sim P(\Theta)
\]

Give every connection a different traffic-defense configuration.

The study integrated its method with WireGuard and reported actual deployment in Mullvad's environment serving many daily users.

Mullvad DAITA v2's dynamic configurations follow a similar direction.

---

38Defense 4:Mixing + DelayOriginal L1489–L1515
# 38. Defense 4:Mixing + Delay

A more fundamental approach than padding:

```text
A ─┐
B ─┼─► MIX
C ─┤
D ─┘
      │
      ├── delay
      ├── reorder
      └── cover traffic
```

to make:

\[
t_{out}
\not\approx
t_{in}+\delta
\]

This actually breaks input / output timing dependency.

---

39LoopixOriginal L1516–L1549
# 39. Loopix

Loopix is representative of this approach.

It uses:

- Poisson mixing
- random delays
- cover traffic
- loop messages

and provides traffic-analysis resistance under a threat model that includes a global network adversary.

The cost is overall message latency on the order of **seconds**, even though this is relatively low for a mixnet.

It is therefore better suited to:

```text
messaging
async communication
mail-like systems
```

Not:

```text
gaming
SSH
remote desktop
interactive low-latency web
```

---

40Defense 5:Multi-user AggregationOriginal L1550–L1595
# 40. Defense 5:Multi-user Aggregation

Another important direction is not adding fake traffic, but:

> Using other users' real traffic as one's own cover traffic.

For example:

```text
Alice ─┐
Bob ───┤
Carol ─┼── Aggregator/Mixer
Dave ──┤
Eve ───┘
```

Inputs:

\[
X_A,X_B,X_C,\ldots
\]

together form:

\[
Y_1,Y_2,Y_3,\ldots
\]

Investigators no longer need only solve:

\[
X_A\leftrightarrow Y_A
\]

but must solve a permutation / assignment problem.

This can increase:

\[
H(Y\mid X)
\]

without requiring all cover traffic to consist of dummy bytes.

---

41Defense 6:Traffic SplittingOriginal L1596–L1641
# 41. Defense 6:Traffic Splitting

Take one flow:

\[
X
\]

and split it into:

\[
X_1,X_2,X_3
\]

across different paths:

```text
             ┌── Path A
Traffic ─────┼── Path B
             └── Path C
```

If the adversary can observe only some paths:

\[
I(X;X_i)
<
I(X;X_1+X_2+X_3)
\]

correlation can be effectively reduced.

For a truly global observer, however:

\[
X_1+X_2+X_3
\]

may still be reaggregated.

The main advantage of multipath is therefore:

> Increasing the cost for partial observers.

---

42Defense 7:Egress TransformationOriginal L1642–L1681
# 42. Defense 7:Egress Transformation

Traditional padding mainly modifies:

```text
Ingress shape
```

But exit connection topology may still remain:

\[
1\ ingress\ flow
\leftrightarrow
1\ egress\ flow
\]

More radical designs can:

```text
Real Flow
   ↓
split / shuffle
   ↓
Virtual Flow A
Virtual Flow B
Virtual Flow C
```

directly breaking:

\[
FlowTopology_{in}
\approx
FlowTopology_{out}
\]

This is a direction more worth researching than packet padding alone.

---

43A new architecture I consider worth researchingOriginal L1682–L1738
# 43. A new architecture I consider worth researching

Combining current research, we can propose:

# Bounded-Latency Multi-Flow Mixing Network

Architecture:

```text
Users

A ─┐
B ─┤
C ─┤
D ─┼── Entry Aggregator
E ─┤
F ─┘
          │
          ▼
   Ephemeral Shaper
          │
          ▼
   Bounded-Delay Mixer
          │
     ┌────┼────┐
     ▼    ▼    ▼
   Path1 Path2 Path3
     │    │    │
     └────┼────┘
          ▼
      Egress Mixer
          │
          ▼
 split / shuffle / merge
          │
          ▼
     Destinations
```

The goal is not to achieve perfect anonymity, but, under:

\[
Latency<L_{max}
\]

\[
Bandwidth<B_{max}
\]

these conditions, to minimize:

\[
I(X;Y)
\]

---

44Layer 1:Ephemeral ShapingOriginal L1739–L1757
# 44. Layer 1:Ephemeral Shaping

For each flow:

\[
\theta_i\sim P(\Theta)
\]

randomly choose:

- padding schedule
- dummy distribution
- packet-size policy
- burst behavior

Avoid creating a fixed defense fingerprint.

---

45Layer 2:Bounded Random DelayOriginal L1758–L1784
# 45. Layer 2:Bounded Random Delay

For each packet / message:

\[
D_i\sim P_D
\]

But:

\[
D_i<D_{max}
\]

For example, different application classes:

```text
Gaming       → very small Dmax
Web          → small Dmax
Streaming    → moderate Dmax
Messaging    → large Dmax
```

This makes the privacy policy QoS-aware.

---

46Layer 3:Multi-user MixingOriginal L1785–L1810
# 46. Layer 3:Multi-user Mixing

Use an aggregation queue:

```text
A packet
B packet
C packet
A packet
D packet
```

reschedule:

```text
C
A
B
D
A
```

Break one-to-one timing mappings within the latency budget.

---

47Layer 4:Real Traffic as CoverOriginal L1811–L1840
# 47. Layer 4:Real Traffic as Cover

Assume the queue already contains:

\[
X_A,X_B,X_C
\]

prioritize using:

\[
X_B,X_C
\]

as A's anonymity cover.

Only when the anonymity deficit is too high:

\[
inject\ dummy
\]

This may substantially reduce:

\[
DummyBandwidth
\]

---

48Layer 5:Egress Flow TransformationOriginal L1841–L1863
# 48. Layer 5:Egress Flow Transformation

At output, also perform:

```text
merge
split
shuffle
connection migration
```

Reduce:

\[
IngressTopology
\leftrightarrow
EgressTopology
\]

its identifiability.

---

49Adaptive Privacy ControllerOriginal L1864–L1905
# 49. Adaptive Privacy Controller

The most valuable use of machine learning is not simply “random padding,” but constrained optimization:

\[
\min_\pi
\left[
\lambda_1L_{privacy}
+
\lambda_2L_{latency}
+
\lambda_3L_{bandwidth}
\right]
\]

subject to:

\[
Latency<L_{QoS}
\]

\[
Bandwidth<B_{budget}
\]

Where:

\[
L_{privacy}
\]

can be jointly approximated by:

- attack ROC
- contrastive matching accuracy
- estimated mutual information
- anonymity-set entropy

.

---

50Defender–Attacker Adversarial TrainingOriginal L1906–L1962
# 50. Defender–Attacker Adversarial Training

Defender:

\[
D_\theta(X)
\]

Outputs a defended trace.

Attacker:

\[
A_\phi(D_\theta(X),Y)
\]

Outputs:

\[
P(match)
\]

The research problem can be written as:

\[
\min_\theta\max_\phi
L_{attack}
\]

Also include:

\[
L_{latency}
\]

and:

\[
L_{bandwidth}
\]

regularization / constraints。

The action space should include more than padding, potentially including:

```text
padding
delay
burst shaping
multiplexing
splitting
reordering
connection migration
```

---

51Classification accuracy should not be the sole security metricOriginal L1963–L2003
# 51. Classification accuracy should not be the sole security metric

For example:

```text
Attack Accuracy:
90% → 20%
```

does not directly prove security.

Because:

> It may merely overfit to one fixed attacker architecture.

A fuller evaluation should include:

\[
TPR
\]

\[
FPR
\]

\[
ROC
\]

\[
Precision
\]

\[
Recall
\]

and candidate population size.

---

52Mutual information is a more worthwhile directionOriginal L2004–L2035
# 52. Mutual information is a more worthwhile direction

What we truly want to reduce is:

\[
I(X;Y)
\]

Or increase:

\[
H(X\mid Y)
\]

Ideally:

\[
P(Y\mid X_A)
\approx
P(Y\mid X_B)
\]

Different ingress flows become increasingly indistinguishable to an observer.

Compared with:

> “Lowering the accuracy of one CNN”

this is closer to an actual security property.

---

53P2P Privacy LabOriginal L2036–L2072
# 53. P2P Privacy Lab

For research, use only lawful data:

```text
random test files
Linux ISO
self-generated datasets
```

Set up:

```text
Client VM
   │
   ▼
VPN
   │
   ├── Tracker VM
   ├── Peer VM
   └── Observer VM
```

and capture traffic at each observation point in:

```text
Ingress
VPN interface
Physical NIC
Tracker
Peer
```

.

---

54Network Isolation Fault InjectionOriginal L2073–L2099
# 54. Network Isolation Fault Injection

Test:

```text
VPN disconnect
VPN server unreachable
route mutation
IPv6 activation
DNS change
Wi-Fi ↔ Ethernet handover
suspend/resume
reboot
client starts before VPN
VPN reconnect
```

Core acceptance criterion:

\[
VPN\downarrow
\Rightarrow
P2PTraffic=0
\]

---

55Traffic-Correlation LabOriginal L2100–L2154
# 55. Traffic-Correlation Lab

Then establish:

```text
Ingress capture:
X

Egress candidates:
Y1
Y2
...
Yn
```

Compare:

### Baseline

```text
WireGuard
```

### Defense A

```text
WireGuard + padding
```

### Defense B

```text
DAITA-like ephemeral shaping
```

### Defense C

```text
bounded delay
```

### Defense D

```text
multi-user aggregation
```

### Defense E

```text
aggregation + shaping + egress transform
```

---

56Recommended core metricsOriginal L2155–L2218
# 56. Recommended core metrics

### Correlation

\[
Corr(X,Y)
\]

### Mutual Information

\[
I(X;Y)
\]

### Attack ROC

\[
TPR(FPR)
\]

Pay particular attention to:

\[
FPR=10^{-3}
\]

\[
10^{-4}
\]

\[
10^{-5}
\]

and other genuinely attribution-sensitive operating regions.

### Latency

\[
P50,P95,P99
\]

### Bandwidth overhead

\[
O_B=
\frac{Bytes_{defended}}
{Bytes_{real}}-1
\]

### Anonymity Set

\[
|S|
\]

Or entropy:

\[
H(U\mid O)
\]

---

57Comparison of defensive capabilitiesOriginal L2219–L2240
# 57. Comparison of defensive capabilities

| Architecture | Real-IP protection | Activity unlinkability | Timing resistance | Latency | Bandwidth cost |
|---|---:|---:|---:|---:|---:|
| BitTorrent encryption | Low | Low | Low | Very low | Very low |
| qBittorrent Anonymous Mode | Low | Low | Low | Very low | Very low |
| VPN | High | Low | Low | Low | Low |
| VPN + bind + firewall | Very high | Low | Low | Low | Low |
| VPN + network namespace | Extremely high | Low | Low | Low | Low |
| Seedbox | Extremely high | Low–medium | Low | Low | Medium |
| Mullvad + Multihop | Extremely high | Medium | Medium–low | Low–medium | Low |
| Mullvad + DAITA | Extremely high | Medium | Medium | Low | Medium |
| Constant-rate shaping | High | Medium–high | High | Medium | Very high |
| Multi-path | High | Medium | Medium | Low–medium | Medium |
| Bounded multi-user mixing | High | High potential | High potential | Medium | Medium |
| Loopix / Mixnet | High | High | High | Seconds | Medium–high |
| I2P-native | High | Higher than clearnet VPN | Medium–high | Medium | Medium |

This table is a relative threat-model-level analysis, not a unified benchmark.

---

58The most important security boundaryOriginal L2241–L2304
# 58. The most important security boundary

The entire problem can be divided into two layers.

## Layer A:Network Deanonymization

Question:

\[
VPN\ Exit
\rightarrow
RealIP?
\]

Main defenses:

```text
VPN
routing isolation
firewall
network namespace
interface binding
```

---

## Layer B:Behavioral / Traffic Attribution

Question:

\[
Activity_A
+
Activity_B
+
Timing
+
Metadata
\rightarrow
SameActor?
\]

Main research areas:

```text
traffic-analysis resistance
mixing
padding
cover traffic
intersection resistance
cross-layer unlinkability
behavioral privacy
```

Perfectly handling the former:

\[
\nRightarrow
\]

the latter is automatically safe.

---

59The right way to position MullvadOriginal L2305–L2357
# 59. The right way to position Mullvad

Mullvad is particularly strong at:

\[
Network\ Identity\ Separation
\]

+

\[
Provider\ Data\ Minimization
\]

+

\[
Traffic\ Fingerprint\ Reduction
\]

In other words:

```text
Shared exit
No activity logs
Numbered account
Multihop
RAM-only
DAITA
```

But it is still:

> **low-latency VPN architecture**

Not:

> traffic-analysis-proof anonymous communication system。

Thus, even if DAITA:

\[
I(X;Y)\downarrow
\]

it is not reasonable to assume:

\[
I(X;Y)=0
\]

---

60Modern investigators actually look for the weakest edgeOriginal L2358–L2410
# 60. Modern investigators actually look for the weakest edge

Assume:

```text
Cryptography       ██████████
VPN routing        ██████████
Real-IP isolation  ██████████
```

But:

```text
Application ID     █████
Timing             ████
Endpoint           ███
Account            ███
Operational error  ██
```

Then a powerful investigator has no reason to:

> Break WireGuard.

Instead, they will:

```text
pivot
```

to:

```text
account
device
time
behavior
provider
metadata
```

Thus:

\[
Security(System)
\approx
\min_i Security(Component_i)
\]

Although not a formal security theorem, this aptly describes the practical attribution attack surface.

---

61Core research conclusionsOriginal L2411–L2486
# 61. Core research conclusions

The study's most important concepts can be condensed into the following points.

### First

\[
Encryption\neq Anonymity
\]

### Second

\[
IP\ Hiding\neq Activity\ Unlinkability
\]

### Third

\[
VPN\neq Anonymous\ Communication\ Network
\]

### Fourth

\[
No\ Logs
\]

from:

\[
Strong\ Traffic\ Analysis\ Resistance
\]

are two different security properties.

### Fifth

The largest fundamental problem in low-latency anonymity is:

\[
Output(t)\approx Input(t-\delta)
\]

Attackers exploit this dependency.

### Sixth

True traffic-analysis defense aims to approach:

\[
P(Y\mid X)\approx P(Y)
\]

In other words:

\[
I(X;Y)\rightarrow0
\]

### Seventh

But:

\[
Strong\ anonymity
+
Low\ latency
+
Low\ bandwidth
\]

have a fundamental trade-off.

---

62The directions most worth further researchOriginal L2487–L2537
# 62. The directions most worth further research

If the research question is:

> **How can end-to-end traffic correlation be reduced as much as possible with minimal bandwidth overhead, within a latency budget that still permits interaction?**

I consider the most promising combination to be neither:

```text
More VPN hops
```

nor:

```text
More encryption
```

It is also about:

\[
\boxed{
Ephemeral\ Shaping
+
Bounded\ Delay
+
Multi-user\ Aggregation
+
Real-Traffic\ Cover
+
Egress\ Transformation
}
\]

and use:

\[
\min
\left(
I(X;Y)
+
\lambda_L Latency
+
\lambda_B Bandwidth
\right)
\]

to establish a formal optimization objective.

---

63Final conclusionOriginal L2538–L2661
# 63. Final conclusion

For ordinary P2P privacy:

\[
\boxed{
VPN
+
Interface\ Binding
+
Firewall
+
Network\ Namespace
}
\]

can already provide very strong protection against:

\[
RealIP\ Leakage
\]

WireGuard's official network-namespace architecture is particularly suitable as a fail-closed engineering boundary.

But if the threat model escalates to:

> An attacker can observe both the ingress and egress of communication and obtain network metadata over the long term.

the problem is no longer a VPN leak.

It is also about:

\[
\boxed{
Traffic\ Correlation
+
Entity\ Resolution
+
Longitudinal\ Attribution
}
\]

This is precisely the core limitation of current low-latency anonymity systems.

Tor's official documentation still explicitly acknowledges that, against an adversary observing both ends of a communication channel simultaneously, the research community knows no practical low-latency architecture that reliably prevents timing / volume correlation.

The genuinely effective strong defenses studied by researchers—such as Poisson mixing, cover traffic, constant-rate transmission, and multi-user mixing—nearly all require:

\[
Latency\uparrow
\]

Or:

\[
Bandwidth\uparrow
\]

or even changes to communication semantics. Loopix is a typical example: it provides traffic-analysis resistance against stronger network observers, but message latency reaches seconds.

Thus the question with real research value for the future is not:

> **“How do we build a completely untraceable VPN?”**

That question is itself ill-defined.

The better question is:

> **“Given latency, bandwidth, and infrastructure budgets, how low can we drive the mutual information between ingress and egress?”**

In other words, study:

\[
\boxed{
\min I(X;Y)
}
\]

subject to:

\[
Latency\le L_{max}
\]

\[
BandwidthOverhead\le B_{max}
\]

This places VPNs, DAITA, Tor, I2P, mixnets, traffic shaping, multi-user aggregation, and future AI-driven traffic defenses within one quantifiable research framework.

---

## Main research sources

**Protocol / Implementation**

- BitTorrent Enhancement Proposals:BEP-5 DHT、BEP-11 PEX、BEP-15 UDP Tracker。
- WireGuard Network Namespace Architecture。
- qBittorrent Anonymous Mode and VPN Interface Binding.
- I2P Threat Model。

**P2P Monitoring / VPN Security**

- Le Blond et al., *Spying the World from Your Laptop*, USENIX LEET 2010。
- Piatek et al., *Why My Printer Received a DMCA Takedown Notice*, HotSec 2008。
- Xue et al., *Bypassing Tunnels*, USENIX Security 2023。

**Anonymous Communication / Traffic Analysis**

- Piotrowska et al., *The Loopix Anonymity System*, USENIX Security 2017。
- Das et al., *Comprehensive Anonymity Trilemma*, PoPETs 2020。
- Holland et al., *DeTorrent*, PoPETs 2024。
- Pulls et al., *Ephemeral Network-Layer Fingerprinting Defenses*, PoPETs 2026。

**Current Deployment / Mullvad**

- Mullvad No-Logging Policy,2026。
- Mullvad Multihop。
- Mullvad DAITA / DAITA v2。
- Mullvad RAM-only VPN infrastructure。

**Case Study**

- CODA:2026 Nyaa First Uploader Arrest / JHA Analytical Tool。

Original SHA-256: 492c5921f53a5ea8e9947dd344a6b1c003f3293eeee20c496ccc75fa65d03d94