Traditional web tracking techniques rely on unique identifiers set in the client-side storage and shared with third-party trackers through network requests. Ideally, this phenomenon may be investigated through the classic lens of information flow control, e.g., by using instrumented browsers with taint tracking support. As a matter of fact though, most web privacy research makes use of simple syntactic matching heuristics that merely look for the presence of (possibly transformed) client-side identifiers within network requests, with no visibility of the JavaScript logic. In this work, we perform a comparative study of these two approaches to web tracking detection. Our investigation shows that taint tracking can expose tracking behavior that remains undetected by syntactic matching heuristics, which suffer from a significant number of false positives and false negatives. However, we also show that taint tracking is not strictly superior to syntactic matching, due to a range of different reasons, including the current limitations of state-of-the-art implementations and the complexity of real-world tracking behavior. Overall, we advocate for a critical reflection on the shortcomings of prominent web tracking detection approaches and we propose useful methodologies to improve current measurement practices.
From syntactic matching to taint tracking and back: a comparative study of web tracking detection techniques / Calzavara, S., Casarin, S., Squarcina, M., Maffei, M.. - 2026:4(2026), pp. 305-320. (PETS 2026 - 26th Privacy Enhancing Technologies Symposium Calgary, Canada 20–25/07/2026) [10.56553/popets-2026-0122].
From syntactic matching to taint tracking and back: a comparative study of web tracking detection techniques
Casarin Samuele
;
2026
Abstract
Traditional web tracking techniques rely on unique identifiers set in the client-side storage and shared with third-party trackers through network requests. Ideally, this phenomenon may be investigated through the classic lens of information flow control, e.g., by using instrumented browsers with taint tracking support. As a matter of fact though, most web privacy research makes use of simple syntactic matching heuristics that merely look for the presence of (possibly transformed) client-side identifiers within network requests, with no visibility of the JavaScript logic. In this work, we perform a comparative study of these two approaches to web tracking detection. Our investigation shows that taint tracking can expose tracking behavior that remains undetected by syntactic matching heuristics, which suffer from a significant number of false positives and false negatives. However, we also show that taint tracking is not strictly superior to syntactic matching, due to a range of different reasons, including the current limitations of state-of-the-art implementations and the complexity of real-world tracking behavior. Overall, we advocate for a critical reflection on the shortcomings of prominent web tracking detection approaches and we propose useful methodologies to improve current measurement practices.| File | Dimensione | Formato | |
|---|---|---|---|
|
popets-2026-0122.pdf
accesso aperto
Descrizione: From Syntactic Matching to Taint Tracking and Back: A Comparative Study of Web Tracking Detection Techniques
Tipologia:
Versione Editoriale (PDF)
Licenza:
Creative commons
Dimensione
744.71 kB
Formato
Adobe PDF
|
744.71 kB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


