Promotor
Prof. Dr. D.C. Thomas
Co-promotores
Dr. F.P.S.M. Ragazzi Dr. W. Soon (University College London)
Financial support
This work was supported by funding from the European Research Council
(ERC) under the European Union’s Horizon 2020 research and innovation
programme (SECURITY VISION, Grant Agreement No. 866535).
Format
This dissertation is available as print, and on the web: https://signals.rubenvandeven.com.
Publishing toolkit: https://git.rubenvandeven.com/security_vision/signals.
Acknowledgements
Writing this dissertation, I have often thought of myself as a loiterer. The loiterer is different from the flâneur, the emblematic figure of 19th-century Paris. The flâneur travels the metropolis’ streets at a snail’s pace (although story has it that some walked their turtles, indicating a velocity just above snail’s pace). Going ever so slow, this observer–participant sucked in their surroundings and showed off. I, however, was but a deer in the headlights.
The loiterer is different also from the wanderer. The wanderer aimlessly travels beyond horizons; the loiterer instead, has an aim yet lingers, unsure which route to take. Idling on an empty page, even my thoughts did not reach the horizon where the sun was setting, rather, they merely ran in circles.
The loiterer sticks to the same street corner for hours on end. The loiterer is a suspect figure, an impostor.
I am grateful then, to all who nudged me into movement.
In the first place, I want to thank my supervisors. By inviting me – artist by training – into political science, Francesco Ragazzi took a big leap of faith. He invested countless hours in getting me up to speed, and even more so in reading, re-reading and reading-yet-again whatever I sent his way. It has been truly remarkable how he fostered alternative infrastructures for making research. And, Francesco’s analogies are unparalleled. I am grateful to Daniel Thomas for joining the leap with us, a bunch of ‘filmmakers’ at the end of the hallway. His trust has been encouraging, and his suggestions invaluable to navigate the PhD trajectory. I am immensely grateful too, that Winnie Soon joined this team, first from Aarhus, then from London. Their outsider’s voice rang close to home, provided a space for sanity and reassurance. They inspired me to ever inquire what ‘coding’ and ‘arts’ can mean for research – and whether we should distinguish these in the first place.
This dissertation would have never happened were it not for the match-making skills of Aymeric Mansoux. Equally important to its realisation were members of the Security Vision research group: Elka Smith, for all your work to keep me on track; Chaeyuen Bae, your work is inspirational, your kindness and support even more so. Special thanks to Ildikó Zonga Plájás, who early on made me enjoy collective writing, packaging its logistics in beautiful neologisms. Thanks also to the Institute of Political Science at Leiden University for providing such a welcoming atmosphere. Josette Daemen, Maria Spirova, and Hannah Kuhn, you deserve a special mention in this regard.
Special thanks to all my collaborators: workshop co-organisers Ana Flamind and Magdalena König, as well as my co-authors Klaudia Klonowska, Sofie van der Maarel, Natalie Welfens and Jasper van der Kist, whose writing is included in this bundle. Our joyful conversations and collective projects worked like medicine for the ‘impostor syndrome’ that is all too common among PhD students.
Key therein too, were all those who engaged with my work at different stages. I would like to thank Anna Leander, Claudia Aradau, Clemens Baier, David Benqué, Jef Huysmans, Lauren B. Wilcox, Mirka Duijn, Nicolas Malevé, Rocco Bellanova, Rune Saugmann Andersen, Soyun Park, Tobias Blanke, transmediale/DARC workshop (Pablo Velasco and Geoff Cox), Next Level Festival (Lex Rütten and Viv Lennert), and several anonymous reviewers. For their help with Perplexity, thanks to Bram Snijders, Arsen Ignatosyan, Stijn Cuijpers; and to Paulus van Dorsten for documenting it. I am indebted to those who shared their expertise with me and who allowed me to enter doors that commonly remain locked.
Crucial to what I do and who I am, are those with whom I interact day-to-day. All Boomtoren inhabitants, and Creative Coding Utrecht affiliates. Fabian van Sluijs, it really is a magical space you have created. Carolien, your work is moving; Werner, thanks for the pun-tjes on the i; Melissa and Joost, you lifted my spirit. Special thanks also to Ward Goes, ever there to hear me out, and stimulate my thinking; we ride together come rain or shine. Cristina Cochior, the way in which you’re both caring and critical is truly inspirational. Sara Orsi, your energy is contagious; Merijn van Moll, your calm reflections are so too. My warmhearted family has always been there. My parents José and Frans, you ever taught and stimulated me, and went out of your way, not seldom at the very last minute.
Most of all, I am grateful to you, Marjo, with whom I hiked many peaks, ridges, and mountain pastures. We dwelt on spectacular views. Our journeys are now joined by the two greatest pranksters: by sending me into a tailspin, they, most of all, set me in motion and helped me put this whole endeavour into perspective.
Summary
Camera surveillance seems like something most of us have little to do with; at most, it is something we merely undergo. This dissertation, which examines how moving bodies are watched and filtered through algorithms, suggests that things are less straightforward.
Surveillance camera operators cannot possibly monitor all the bodies moving through a city. They have to continuously decide whom to watch, and for how long. Such targeting practices are increasingly supported by computer vision systems, which analyse movement to distinguish ‘innocuous’ passers-by from those relevant to the operator. But how do these systems attribute suspicion? And what does it even mean for an algorithm to ‘perceive’ movement?
To analyse the underlying mechanisms, this dissertation introduces re-modulation. Similar to reverse engineering and re-modelling, re-modulation unravels a system’s components by remaking them. Re-modulation is, however, a form of multimodal, artistic research that highlights the process of adjusting, which continuously shifts the relationship between researcher and system.
Drawing on this technical engagement, this dissertation borrows concepts from signal processing and information theory to reformulate the role of the surveilled subject. In research on algorithmic security – whether real-time or pre-emptive – ‘suspicion’ typically forms an implicit binary category: a subject is labelled either “suspect” or “not suspect”, as with the terrorist, smuggler, or migrant. The signal subject, instead, forgoes this fixed threat category. When an operator selects a moving body as a target – turning the camera in their direction and zooming in – its relevance is determined by a continuous justification of the observed movement. A relevant target disrupts anticipated normality. Suspicion, then, is registered as ‘surprise’: the operator’s failure to predict.
The logics of anticipation and surprise that enable the algorithmic ‘perception’ of movement, this dissertation argues, shift the conventional distinction between surveiller and surveilled. The surveilled does not merely undergo observation. As users of public space unselfconsciously go about their daily routines, they form the background of normality against which others stand out. Thus, the majority of passers-by are not impaired by observation, but granted agency as they shape how others come to be judged.
Samenvatting
Cameratoezicht, of -surveillance, lijkt iets waar de meesten van ons niets mee te maken hebben; hooguit hebben we het te ondergaan. Dit proefschrift, dat onderzoekt hoe bewegende lichamen worden bekeken en gefilterd met algoritmes, suggereert dat de zaken complexer liggen.
Camera-observanten kunnen onmogelijk alle lichamen monitoren die zich tegelijkertijd door een stad bewegen. Voortdurend beslissen ze wie ze in de gaten houden, en voor hoe lang. Zulke selectiepraktijken worden steeds vaker ondersteund door computer vision-systemen, die bewegingen analyseren om ‘onschuldige’ voorbijgangers te onderscheiden van personen die relevant zijn voor de observant. Maar hoe worden afwijkende gedragingen gevonden? En wat gebeurt er eigenlijk wanneer een algoritme ‘beweging’ waarneemt?
Om de onderliggende mechanismen te analyseren, introduceert dit proefschrift re-modulatie. Vergelijkbaar met reverse engineering en re-modelling, worden de componenten van een systeem ontrafeld door het na te maken. Re-modulatie is een vorm van multimodaal, artistiek onderzoek, dat de voortdurend veranderende relatie tussen onderzoeker en systeem in ogenschouw neemt.
Door aan deze beschouwing concepten te ontlenen uit de signaalverwerking en informatietheorie, herformuleert dit proefschrift de rol van het geobserveerde subject. In onderzoek naar algoritmische beveiliging – ‘real-time’ of preventief– vormt ‘verdacht’ gedrag doorgaans impliciet een binaire categorie: een subject wordt gelabeld als “verdacht” of “niet verdacht”, zoals de terrorist, smokkelaar of migrant. Bij het ‘signaal-subject’ wordt geen vastgestelde dreigingscategorie toegekend. Wanneer een observant een bewegend lichaam als doel kiest — door de camera te richten en in te zoomen — bepaalt een voortdurend gezochte verklaring voor de waargenomen beweging de relevantie van een doel. Een relevant doel onderbreekt de verwachte normaliteit. ‘Afwijking’ wordt dus geconstateerd als ‘verrassing’: het onvermogen van de observant om beweging te voorspellen.
Vanuit het oogpunt van ‘verrassing’, zo betoogt dit proefschrift, verandert het gangbare onderscheid tussen bewaker en bewaakte. De bewaakte ondergaat de observatie niet louter passief. Integendeel: terwijl gebruikers van de publieke ruimte onbevangen hun dagelijkse routines uitvoeren, vormen zij de achtergrond van normaliteit waartegen anderen afsteken. Zo wordt de meerderheid van de voorbijgangers niet belemmerd door observatie, maar beïnvloedt zij de beoordeling van anderen.
1 “I hear nothing but static”
Introduction
It was a sunny Sunday morning; I was up early. Brushing my teeth, I absent-mindedly looked out of the bathroom window. The street below was still but for a single human figure, walking, a bag hanging over his shoulder. I saw the man, but I did not notice him; plenty people walk by the house every day. Going by the bag, he was likely going to do a Sunday morning workout in the gym or the nearby forest. The figure turned into my neighbours’ driveway across the street, and tried to open the gate. It was locked. Again, I did not consciously register – after all, he might be a friend of theirs. When figure turned, I half expected him to walk up to their front door and try the doorbell. He did not. Instead, he walked up the driveway of the next-door neighbours. This was unexpected, only now I noticed him; he had my curiosity. He clearly was not delivering mail, or newspapers, but he might have mixed up the houses the first time around. He turned around again, skipped the next driveway, and walked into the garden of yet another neighbouring house. Now, the figure had my attention.
Well into my research, while reflecting upon the material I had collected – a pile of fieldwork notes, code fragments, schematics, and literature references – I remembered this situation. Out of all the people walking by my house, what was it that made this figure “stand out”? It could not have been his sole presence on an empty street, for if the figure had followed the pavement, I would have hardly noticed him. I vividly recalled how my attention gradually shifted away from my daily routine toward the figure in the street. All of a sudden, I had found myself taking the role of surveillor, assessing someone unknown to me. Time and again, the figure violated the movements I unconsciously anticipated; thereby defying a mundane explanation for his presence and intentions. Ultimately, the dissertation in front of you concerns this process: for what happens in the transition from being observed to standing out?
This question has become all the more pertinent over the past years. Since cameras have been introduced into public space only a few decades ago, they have become ubiquitous1. From centralised hubs, camera operators observe visitors of city centres, shopping malls and train stations. They are on the lookout for disturbances and indicators of future peril. However, the increasing number of cameras produce so much footage that municipalities do not have enough personnel available to constantly monitor all video feeds. It is not uncommon that one person has to take care of more than a hundred simultaneous streams. In recent years these observer rooms have seen the introduction of algorithmic techniques that are to guide the operator’s eyes by singling out particular behaviours. By highlighting specific camera feeds, these systems set observation priorities for operators.
These technologies augment conventional camera surveillance with so called ‘computer vision’ techniques that automate the analysis of video. Such systems are promised to detect all kinds of anomalous situations: pickpockets (Bouma et al., 2014), violence (Oddity.ai, 2026), or supposed burglars (Hamada, 2020), crowd congestion and people exhibiting suicidal behaviour (Kreissl et al., 2014: 170; Lyon, 2003: 273). Many of these systems are in their infancy, used only in academic contexts, but some are being trailed in public space, and a few have already been used in day-to-day municipal surveillance operations (Hulsen, 2024).
In a sense, these systems are to automate the task of so-called ‘spotters’: police officers or other security personnel who use visual cues to assess suspicious behaviour. Recognising such behaviours makes rapid intervention possible whenever an incident occurs. An example that put spotters in the spotlights, making it into Dutch prime-time news, occurred in October 2018, when Jawed S. stabbed two tourists with a knife in Amsterdam Central Station. Prior to the event, S’s behaviour had already been signalled by police present at the station, and within nine seconds after his first stab, police officers apprehended him. ‘We point out to [our colleagues patrolling] which people stand out: they deviate from their walking trajectory, or they linger around. These people are of interest to us.’ (NOS nieuws, 2018) Spotters thus examine a subject’s movements, and guide the attention of law enforcement. Much like human spotters, the systems central to this thesis serve to distinguish the irrelevant – innocent even – passerby from those ‘of interest’ to surveillance: protecting one from the other.
Set at recognising behaviours, these systems serve a purpose different from oft-cited biometric technologies. Biometry has no single definition, but commonly denotes practices in which the appearance of the body serves as evidence for a claim – for example of identity (Maguire et al., 2018; Pugliese, 2010). Think of fingerprinting technologies used at passport checks; or of camera based automatic facial recognition that is deployed at airports, or at shopping malls, to detect people on a blacklist, or at elderly homes to prevent patients with dementia leaving the building while keeping the door open for others. Of course, behaviour recognition technologies also analyse the human body as it appears on images, and thus in a literal sense, they could be termed ‘biometric’. Yet these technologies rely less on phenotypic classification, instead examining the body’s movements.
But how does the algorithmic analysis of movement produce surveillance targets? That is, how does the algorithmic analysis of movement filter innocuous passersby from those of interest? How are these video based systems different from those that rely on still images to identify individuals, and how does that affect the ways in which suspicion is attributed to individuals? And underlying all these questions, what does it mean to ‘perceive’ movement algorithmically anyway?
This dissertation is thus set at understanding how security devices inform the sensemaking of security practitioners. To that end, the different elements of this dissertation tease out the techniques fundamental to the algorithmic perception of movement. It traces these algorithmic surveillance systems, often hailed as ground-breaking ‘Artificial Intelligence’ technologies, to signal processing techniques that originate from the 1960s2. This dissertation draws on concepts from signal processing and information theory to develop an understanding of security and its moving subject.
Tracing these algorithmic systems is not just a technical, analytical exercise. Years into writing this dissertation, I was traversing Utrecht Central Station, mulling over a question that I got confronted with yet again after presenting my research: ‘where is the politics?’
Clearly, something had thus far evaded my descriptions. I told myself I was caught between a technical fascination with these devices and a reluctance to default to the common critiques of algorithmic power, which only partly fitted the case at hand (cf. Galič et al., 2017). Either stance risked overlooking the intricacies of what is at stake in the use of these technologies, and would end up reifying promotional narratives of ‘Artificial Intelligence’ and technological advancement (Aradau and Blanke, 2022: 9; Hunger, 2023).
All of a sudden I held still and looked around the station. I had found myself routinely taking my pathway to the train, just like anyone else in the station. But after years of taking in critical literature, was I not supposed be conscious of the cameras hanging there? It had only been a week since I had visited a municipal CCTV operation room. Was I not to feel the piercing eyes of operators sitting behind their desk? To be sure, I minded be watched – when I thought about it – but I, I realised with some sense of shock, did not care.
Moral crisis ensued. Slowly it dawned on me that rather than careless, I had ever been care free: I check all ‘seven checkmarks’ (Luyendijk, 2022): I am white, born in the Netherlands, affluent, male, heterosexual, finished higher education. Hardly ever did I have to worry about encounters with law enforcement. How was I ever to develop an honest politics of algorithmic surveillance as a device of power from such a privileged position? Where did this indifference leave me with a critique of algorithmic surveillance?
To gain purchase on a politics of these security technologies, I this dissertation shows, we need not only to unpack the machines, nor analyse the discourse: now and then we have to hold pause, and reflect on how these devices make us feel.
In this dissertation, technological description serves not as an end in itself, but facilitates an affective and reflexive practice. Cross reading such descriptions with political analysis leads me to stake two claims regarding security practices that analyse moving bodies. First, an examination of the algorithmic analysis of movement forces us to think of suspicion not as a matter of categorisation – in which a subject either is, or is not, suspect (cf. Dijstelbloem and Broeders, 2014; Lyon, 2003: 276) – but as a process. The process of targeting a moving body is less about establishing suspicion in light of a pre-established threat (Massumi, 2010), than it is about justifying observations in terms of normalcy.
Second, we need to reformulate the relation between a surveilled general public and those who come to stand out as targets. Such a reformulation of relations is necessary, I claim, for the critique of surveillance as a practice that harms the surveilled subject – infringing privacy, risking misclassification – has become stale. Suspicion is not a damoclessian sword that hangs above the head of every subject (cf. Amoore and De Goede, 2008: 180; Bigo, 2014: 219). Rather, most subjects are simply safe from becoming a target. Examining the operations by which the attention of the operator is guided, and dissecting these through logics of signal and noise which are so central to algorithmic processing, I posit that in fact the subject’s acts inform the assessment of others. Thereby, the subject conforming to its routine movements, rather than harmed, is complacent to surveillance.
Before examining these claims, in this introduction, I will first touch upon the bodies of literature that undergird this research, and briefly discuss the notions of security, algorithms and movement that inform my work. Subsequently, I will trace a red thread through the different elements of the dissertation, by discussing how they engage with the central questions, both conceptually and methodologically. In this dissertation’s conclusion, I will reformulate the surveillance of movement in terms of signal processing to I explore the implications of these security technologies for how we think about surveillance, and plot avenues for further research.
Algorithmic security devices
This dissertation thus revolves around the operations by which moving subjects are surveilled, and examines the algorithms that make some stand out from the rest. The research draws on a substantive body of work that looks at the intersection of security, human movement and technology. Two bodies of literature have been particularly informative.
First, the disciplines of critical security studies and surveillance studies provide crucial reflections on the role of technology in security. Both fields coalesce around the dilemma that ‘security’ is not some kind of objective value, but a practice in which establishing something as a security concern is a political act that serves to legimitise interventive actions. While these actions produce securities for some, they simultaneously create insecurities for others (Ashley and Walker, 1990; Bigo, 1996; c.a.s.e. collective, 2006; Wæver, 1998). Technology automating decision making, normalises given political orders and justifies the practices in which they play part.
Algorithms that examine movements should thus not be taken as neutral devices that merely support existing security forces in their task; for the security task is not distinct from the technology that enables its enforcement. Rather, algorithmic operations are constituted in, and constitutive of a broader network of power relations, that shape everyday ‘security’ practices. At stake in such critical examinations of security is how power and agency is configured and reconfigured as the instrumentarium of security changes (e.g. Amicelle et al., 2015; Hoijtink and Leese, 2019a).
The entanglement of security concerns with technologies and techniques of governance does not mean that security practices are ever at the whim of technological change. To delve into the effects of technology on power, this research draws on a second body of literature, which coalesces around the field of media studies and its subfields software studies (Fuller, 2008; Kittler, 1995), and critical code studies (Marino, 2006). This branch of research grants particular attention to the material conditions by which meaning is produced, circulated and consumed. The conversion steps that happen between the hardware that records phenomena, and the (human) perception that is to interpret it – so highlights this literature – are at the same time logical and symbolic, routine and rhythmic, functional and affective. Algorithmic logics are thus inseparable from the imaginations and imaginaries around what code and coding is and should be doing: “the wider the gap […] the wilder the imagination.” (Cramer, 2005; see also Cox and McLean, 2013; Ruckenstein, 2023) Tracing technology’s exact operations and placing them in lineages (Huhtamo and Parikka, 2011) breaks the ‘myth of the new’ that fuels the imaginaries that surround them. Descriptions of the mechanical thereby become the entry point to discuss how contemporary security practices sense and perceive threats.
Neither body of literature thus takes its object of study – security (c.a.s.e. collective, 2006), algorithms (Gillespie, 2014; Seaver, 2017) – for granted, examining instead how they are configured in collective practices. Algorithmic security practices are thereby understood to bring together of all kinds of entities, such as people, technologies, procedures, and regulations. What is at stake then in algorithmic camera surveillance, and how it exercises power over moving bodies, are the “mico-physics” of forces that make these objects rhizomatically connect and ‘work’ together as a single entity, and which make movement intelligible as a governable phenomenon (e.g. Huysmans, 2023; Hoijtink and Leese, 2019b).
Data flows: surveilling movement
Across these different fields of research, algorithmic surveillance is often understood to govern and produce movements in terms of ‘flows’. Computer technologies accumulate all kinds of records: social media profiles, payment transactions, travel logs. Using ‘data mining’ techniques, these records are amalgamated into profiles, or ‘data doubles’ (Raley, 2013a; see also Muller, 2008: 128; de Vries, 2010; Hildebrandt, 2008). These data doubles, often adhering to strict data formats, both represent, yet are distinct from the corporeal moving body (Goriunova, 2019). These digital subjects are bodies of data that ‘circulate’ through databases as ‘discrete flows’ of information (Haggerty and Ericson, 2000). Data bodies, transferred in ‘streams’, move much faster and more freely than a physical body ever can. Therefore, each flow is subjected to a different security regime that governs its movements (Bigo, 2014).
As an abstraction for movement the ‘flow’ has a certain appeal. Flow is reminiscent of liquids, or from a materialist perspective, liquidity (Bauman, 2013; Lyon, 2010). When data flows, it creates ‘behavioral surplus’ that translates to shareholder value (Zuboff, 2019; see also Aradau and Blanke, 2022). Flows can become waves that crash into a nation’s shore, or a deluge that washes away everything one holds dear. Flows thus invite control: they can be channelled, funnelled, put under pressure, boiled or frozen. Flows fill ‘data lakes’, reservoirs of unstructured in-formation. A well-placed dam can harness the flows’ potential energy, and convert it into power. Most importantly, flows homogenise: they unify moving matter (see also Nail, 2022). While their composition can change with time, any distinction between individual bodies is washed away, dissolved into somewhat stable trajectories. Operating on ‘flows’, security practices massify individual bodies into populations, crowds (Nishiyama, 2018), and clusters (Isin and Ruppert, 2020), the movements of which can be monitored and directed as a whole.
The metaphor of the ‘flow’, however, fails for the research presented here, for it is not concerned with power over populations, but with the ways in which people come to stand out. Opposite to the amalgamation of moving bodies and their digital representation into flows, it is concerned with their differentiation. Therefore, across this thesis’ various sections, I explore an alternative terminology with which to describe how movement becomes a governable phenomenon.
This research thereby also satiates a personal curiosity: for what is a data flow anyhow, and how is a moving body ‘cast’ into one? Having worked as programmer, I have experience managing heaps of data and transferring millions of database records. Coming to critical literature on technology and security, I have found the proliferation of the term ‘flow’ somewhat curious: it is reminiscent of technical terminology, implying an explanation (‘but of course, data flows moves’), yet at the same time the term is rather evasive of specificities, for what it is precisely that moves, and how it is made to do so? If anything, the flow can be related to the ‘data stream’, the perceived homogeneity of which is merely an approximation that emerges from a discontinuous process (Soon, 2019; see also Ernst, 2011). There thus is a fissure in the continuity of movement, which becomes visible only when zooming in on the minute operations at play. Therefore, this dissertation develops its analysis at the interface of technical detail and politics, and explores new methods to do so.
Algorithmic techniques
“Technicalities themselves, in their most intimate details, are technically underdetermined. They depend on social matters: practicalities, contingencies, power plays, traditions. Thus, technicalities should not be left to professionals alone. They affect us all, for they involve our ways of living. But this does not mean that they are not also technicalities.” (Mol, 2002: 171)
Nitty-gritty specificities matter, so stresses philosopher of science Annemarie Mol, for a system’s technicalities shape its political effects. In the field of critical security studies, attention to the minutiae of technological devices has recently gained prominence. In particular in research that draws on science and technology studies and media studies, technicalities are not only analysed at the human-machine interface – by means of participant observation, legal analysis or interviews – but also by examining machine-machine interactions and how these are governed (e.g. Aradau and Blanke, 2022; Amoore, 2020; Bellanova and Glouftsios, 2022; Planqué-van Hardeveld, 2023; Pötzsch, 2015).
Such accounts of machine-machine interactions offer several benefits. Most importantly, it avoids a common pitfall in critical analysis, that still too often defaults to a monolytic narrative of ‘the algorithm’, which is portrayed as a “black box”. This algorithm-as-black-box narrative produces two diametrically opposite analytical responses. The first is that such a box can be understood by simply unwrapping them, and glimpsing what is inside (cf. Pasquale, 2015). But this is a fallacy. As many authors have noted, one can only be disappointed by what is inside: algorithms are like infinite Matryoshka dolls, they can be unwrapped, only to reveal another black box (Introna, 2013; Leese, 2014). The conclusions that these authors then arrive at is that the algorithmic ‘black box’ is fundamentally unknowable: the sheer scale at which algorithmic assemblages operate makes them too vast to grasp. Instead, these authors take their cue from early cybernetics, suggesting all one can do is test the box: by providing an input and examining how it responds a researcher can empirically derive a mental model for its functioning (Bellanova and Fuster, 2019; Bucher, 2018: 44).
Relational understandings of software, however, provide a middle ground between these two positions. Underscoring this modularity, Rieder (2020) suggests we can get to know algorithmic practices by examining the algorithmic techniques that amalgamate in complex assemblages (see also Burrell, 2016; Mackenzie, 2017). An algorithmic technique is similar to the traditional definition of an ‘algorithm’ as a more-or-less formalised routine, such as Bubblesort (see also Jaton, 2017). Yet analytically, ‘algorithmic technique’ stresses the socio-political embeddedness of computational operations that “are at the same time material blocks of technicity, units of knowledge, vocabularies for expression in the medium of function, and constitutive elements of developers’ technical imaginaries.” (Rieder, 2020: 17) Techniques are thus not so much formalised structures, but emerge by their material-semiotical circulation: as written code libraries, as documentation thereof, as narratives, as solutions, as technical debt (see also Seaver, 2017). They are “ideas in matter” (Portanova, 2013: 8) that paint the horizons of possibility. Algorithmic techniques can be composed to form ever more complex systems, and decomposed to technographically trace the operations at play (Bucher, 2018; see also Latour, 2005). The algorithmic system, as an assemblage of techniques, is a grey, rather than black box. Even though we cannot keep infinitely “opening up” the algorithm, and examine its components – we are not completely oblivious to its internals either.
Re-modulation
The color of the box notwithstanding, the technicalities of an algorithmic surveillance assemblages are fractically complex, as the technological ‘stack’ is nearly infinitely layered. This is further complicated by contemporary software that makes such assemblages ever more unstable: any ‘over-the-air update’ can alter the code and data that are executed. Especially in their trial phase, as perpetual beta, surveillance technologies are moving targets of analysis (Mackenzie, 2006). How then, when studying this network of relations in which matter and meaning is co-constituted (see also KM Barad, 2007), can one cut into the assemblage and bracket the analysis?
To navigate this onto-epistemological concern, I propose a methodological stance that throughout this dissertation takes different guises, but which I frame under the conceptual umbrella of re-modulation. With re-modulation I draw on methods of art and artistic-research as well as methods developed in software studies such as exploratory programming (Montfort, 2016) that encourage the researcher to work with the concrete lines of code that make the systems they analyse tick. Most explicitly, ‘re-modulation’ is a play on ‘re-modelling’ (Miyazaki, 2020), a research approach that intervenes in the schematics of an algorithmic system. Re-modelling implies a hands-on re-implementation of the systems analysed. Such an implementation “is not merely about copying a process” but “about generating new maps by slightly changing some parameters and relations entangled with such a process.” (Miyazaki, 2020: 242) Similar to re-modelling, a re-modulation reconstructs an algorithmic system, which requires the (re-)maker to take it apart into modules, not once, but time-and-again, to an ever finer degree (Soon and Velasco, 2023; see also Pelizza and van Rossum, 2021). Here, the intention behind coding is thus not to (re)construct a box so that I, the researcher, can examine its output, as if the software’s politics become visible when its code is executed (e.g. Mackenzie, 2017; see also Fuller, 2008). Rather, recoding a system, writing the instructions, the researcher gains some understanding of the operations at play, and in the process heuristically runs into whatever needs to be understood for the system to ‘work’.
Re-modulation, as a making method, probes the aesthetics of technological security devices. Therein, it shares similarities with visual analyses of security, which, over the past decade, have opened up security research to non-text (written or spoken) registers (Bleiker, 2018; Galai, 2023; Vuori and Andersen, 2018). However, re-modulation is explicitly a multi-modal method by which research is done and disseminated across different registers (see also Austin and Leander, 2021; Möller et al., 2022; Ragazzi, 2025). The visual, the aesthetic, is thus not external to, but intrinsic to the research.
A technographic methodology of re-modulation forefronts a reflexive stance. Re-modulating a surveillance system for the purpose of political analysis is not a mere reverse engineering; it does not aspire a 1:1 reproduction. In the process of employing the software in a research context, its configuration changes, as its new context is radically different from the surveillance context for which it was devised. Modulation – a term borrowed from the field of signal processing that is so central to this dissertation – shifts the parameters of a wave such that it carries new information. Thus rather than seeking to fix, critique, or expose, re-modulation is a continuous adjustment. Each parameter change reconfigures the researcher in relation to its object (cf. KM Barad, 2007). So even though, as researchers, we tend to frame our work as a coherent whole, re-modulation is more of a stumbling-into-things; a “deliberate articulation of […] unfinished material thinking.” (Borgdorff, 2012: 71; see also Wesseling, 2016). Re-modulation, in that sense, is a coming to terms with the unease of that one, recurring, question – ‘where’s the politics?’ – by engaging with the researcher’s positionality. Thereby, re-modulation is an artistic approach that provides space to reflect upon one’s own discomfort. It purposefully – and occasionally compulsively – steers away from finite resolution (Goatley, 2019: 27), it refuses to pin down, and postpones judgement. In a re-modulation, a politics of technology is ever tangled up with a politics of research.
Normality machines
The analysis presented in this dissertation has been developed across different registers: written text, computer code, an interactive installation. Each of these comes with its own modalities of circulation. Echoing the fragmented character of the research and its dissemination, this dissertation is not a monograph, but bundles four texts, which have been published (or are in the process of being published), in different outlets. Nevertheless, all four texts engage, explicitly or implicitly, with the dissertation’s central question: how do surveilled subjects come to stand out by means of computer vision techniques?
1. Sensing security
The first paper in this bundle starts from the observation that across fields as diverse as defence, police and migration control, algorithmic data processing is collapsed under a practice of ‘sensing’. How can we as researchers of security relate to such a description of technology? The paper documents a year of collective discussion between me and four other scholars, who share a concern with the non-neutral ways of ‘sensing’ in, and ‘making sense’ of security practices.
Speaking of sensing, my co-authors and I argue, implicitly mobilises a belief that these technologies merely augment the human sensorium, improving on its sensitivity and reach. However, sensors are not mere prosthetics external to the human body, providing access to new sources of information, but they “are changing the subjects of experience as well” (Gabrys, 2016: 22, see also p.65). As such, we argue, the composition of sensing technologies matters for how threats are perceived and interpreted, while shaping the security practices of which they are part.
The sensor itself is not just a measurement instrument, a mere input
point of data, but produces its signal by means of computational
techniques. What is called a sensor can often be unpacked as
camera+computer. Such computational sensors enact a
sequence of ‘transductions’ (Mackenzie, 2002), each transforming,
converting and interpreting the sensor input to produce an output
signal. At each step, these transductions articulate and rearticulate
measured realities (see also Steyerl,
2014). Sensing practices deploying
camera+computer=sensors reconfigure relations of power
between surveiller and surveilled.
At stake in ‘vision’ of ‘computer vision’, is thus not the relation to the human eye, but the onto-epistemological formations of data-processing such technologies gives rise to. Hence, the collective paper advocates, critical study of security should describe the cascades of transformations taking place to dissociate sensing practices from the human senses and open the monolithic sensor up to alternative compositions.
2. Diagramming algorithmic politics
The second paper this dissertation bundles, turns to sociotechnical practices in which ‘computer vision’ and ‘security’ feed into each other, to further inquire how such technologies enable practitioners to see, sense and surveill. The paper grew out of a desire to map the use of computer vision technologies in security practices, in order to flesh out relevant research questions. But how is one map algorithmic practices? Early on in our research, my co-authors and I attempted such a mapping by neatly categorising the practices we encountered – databases, companies, algorithms, etc (see Ragazzi et al., 2021). However, categorising complex algorithmic devices proved difficult (for example, databases are not singular systems, they run ‘algorithms’; which algorithms then would deserve their own entry, which do not?). We kept changing the schema of the collected data: refining categories, adding attributes and relations. Yet, we felt ever-more conflicted: were we not reproducing the rigid systems of classification our research intended to inquire? Were we not mobilising the exact same aesthetics of data visualisation that we had become so sceptical of? (see Fuller and Weizman, 2021; Drucker, 2011; Plájás et al., 2020 documents our conversations; see also Ragazzi, 2023). In retrospect, the questions we had then – the reluctance to pin down practices into marked categories, and the unwillingness to disentangle our own position from the aesthetics of research – touch upon this thesis’ methodological proposition: re-modulation inquires the tools by which we get to know algorithms.
The resulting paper proposes an exploratory methodological device, a time-based diagramming tool. Diagramming is an open-ended, associative and processual approach that complements interviews with a real-time drawing of diagrams3. The paper maps out how the boundaries of algorithmic systems are drawn and negotiated by practitioners, and how entities solidify and stabilise as they circulate between sites of development and deployment. By examining the drawings of the interviews, the paper sketches out a rhizome of interrelated processes and politics, through which my co-authors and me draw five pathways, each indicative of shifting politics of algorithmic security.
Two of the shifts highlighted in this section are of particular relevance for the computational assessment of movement, and the filtering of suspicious subjects. The first concerns a shift in algorithmic perception, from what one might call a photographic vision to a cinematic vision. A photographic vision is about knowing the individual, it relies on the inscription of supposed ‘indelible’ truths in supposedly static biological features of individuals (faces, fingerprints, retinas). Of course, such a photographic logic of algorithmic perception can be applied to with sequences of snapshots that are taken over a period of time. However, as increases in processing power significantly decrease the interval between individual sensor readings, a different interpretative logic becomes possible. This cinematic vision operationalises the variability of life, thereby constructing suspicion as time-based, ephemeral and changing.
A second shift highlighted in this section makes apparent how an algorithm’s failure to make the right prediction is not a defect of algorithmic systems, but a central characteristic. In ‘machine learning’, the most prominent technique for contemporary computer vision, a model converges by means of an error metric. This metric, due to the complexity and ambiguity of life, can never reach zero. If anything, an error-free model is suspect, considered ‘overfitted’ on the training data, and as failed to generalise on the patterns that it had been presented with. Consequently, in computational sense making, reality can never be known for certain. The inevitable presence of errors in the predictions of algorithmic systems means security practitioners can not blindly rely on these results. This uncertainty influences how these devices are integrated in security practices.
3. Prediction of the now
The third text of this dissertation builds on both insights – the cinematic algorithmic vision, and the inevitable uncertainty of prediction – as it addresses a deceptively simple question: what is movement in algorithmic terms? The paper turns to DeepSORT, a prominent algorithmic technique for ‘multi-object tracking’ that transduces immobile, frame-by-frame, detections into moving subjects. By means of a re-modulation of the tracker, the paper technographically traces the logics at play, and contends there is no observation of movement without its assessment.
By examining concepts from signal processing and information theory that movement tracking relies on, I suggest these technologies require rethinking the subject of security. The political effects of algorithmic movement tracking cannot be accounted for by existing debates on technology in security. In these debates, the political effects of technology are either analysed within a linear temporal framework of immediacy – in which new technologies make subjects move faster or slower (e.g. Leese and Pollozek, 2023; Walters, 2017) – or they pertain to the calculative techniques by which the future is anticipated in terms of potential harm (e.g. Amoore and De Goede, 2008; Aradau et al., 2008; Lyon and Wood, 2021; Massumi, 2010). However, algorithmically accounting for movement not only increases the velocity of a still subject, but places its perception in a different temporal logic.
The section upholds that with the algorithmic tracking of movement, the ‘now’ is predictively enacted through an anticipatory model of normality. With the prediction of the now, the subject under surveillance is no longer assessed in terms of its similarity to a ‘risky other’, or in terms of potential future harm. Instead the other emerges by logics of anticipation and surprise in the temporal variability between algorithmic prediction and measurement. This section draws on formulas from signal processing, which are to resolve the fundamental variability and uncertainty of measuring reality, to formulate an alternative role for the subject under surveillance. It concludes that, when the subject is governed by means of a model of normality, it becomes complicit to the assessment of others.
4. Perplexity
This reconfiguring of the moving subject of surveillance is further developed in an interactive art installation that is part of this dissertation, and which is complemented by short essay: Perplexity. Perplexity re-modulates an algorithmic anomaly detector. Such algorithmic systems track bodies as they move through space, recurrently predicting its future movements. Suspicion emerges when a subject deviates from the predicted trajectory.
Consisting of LiDAR sensors, a computer, and laser projectors, the installation re-appropriates these generative surveillance algorithms. Like an anomaly tracker, Perplexity tracks people passing through space, and captures their trajectories. The captured tracks are used to continuously train a model to forecast the movements of others. In this re-modulation, the forecasted trajectories are drawn right in front of the passerby’s feet as pathways of light. Thereby, the installation creates an interaction between the pedestrian and the predictive algorithm. The participant is not only interpreting the object, but its movements inform the unfolding of the piece for whoever passes the space at a later time.
Re-modulating an anomaly detector, and reflecting on my own position towards surveillance as I traverse public space, this section brings to bear the role of the surveilled as co-constutitive of surveillance. Descriptions of surveillance, algorithmic or not, tend to describe a one-way flow of information from the surveilled to the surveillor. Such descriptions render the distinction between surveiller and surveilled as a binary. Yet thinking through security practices in terms of signal/noise, and anticipation/surprise, upsets this binary. Security practices govern movement by measuring conformity (see also Huysmans, 2023), and maintaining normality. When we consider the production of that normality, which is so explicitly modelled in algorithmic surveillance, it becomes apparent that any observed subject contributes to the predictive mechanics by which the other’s movements are anticipated and thus assessed.
To an observer, most of our movements are static4. This noise is irrelevant to their task. Suspicion, then, is established over time, by gradually filtering a signal from everything and everyone else in the ether. Consequently, we should not discount the role the surveilled subject plays in the surveillance of others. For standing out is a standing out against others, and thus the other stands out against you.
2 Sensing Security
Collective Discussion
Authors: Klaudia Klonowska, Sofie van der Maarel, Jasper van der Kist, Natalie Welfens, Ruben van de Ven
This text is published International Political Sociology (2025). 10.1093/ips/olaf044
Note on inclusion in the dissertation: This article is the outcome of a year-long collaborative scholarly dialogue. Each author provided their own sectional contributions, and the authorship of the introductory and concluding frameworks was shared. However, in accordance with PhD regulations regarding co-authored works, I assume primary responsibility for the integrative stage of the article, drafting the concluding arguments in which the various individual contributions were integrated into a coherent whole.
The fields of International Political Sociology (IPS) and Critical Security Studies (CSS) have examined the various ways in which people and objects are interpreted and perceived as security threats. This scholarship has recently gained new momentum due to the increasing use of sensor technologies, which are changing how security is understood and practiced. This collective discussion piece brings together five scholars who share a concern with the non-neutral ways of “sensing” and “making sense” of security. The discussion advances the research agenda by examining the mediated practices of rendering security perceptible and actionable, and examining its political implications. Illustrated through empirical cases, we make three contributions to the discussion about sensing (in)security. First, we examine how sensing technologies are understood to imitate or extend human perception. Second, we consider how technological mediation simultaneously reduces and amplifies sensing and sense-making. Third, we analyze how these processes are constitutive of relations of power.
Introduction
Scholars working in the fields of International Political Sociology (IPS) and Critical Security Studies (CSS) have long studied the various ways in which persons and objects are interpreted and perceived as security threats. This scholarship has recently gained new momentum due to the increasing use of sensor technologies, which are changing how security is understood and practiced. This collective discussion piece brings together five scholars who share a concern with the non-neutral ways of “sensing” and “making sense” of security; how the variable worlds of crime and policing, borders and migration, military and warfare are made perceptible in particular ways so that they can be acted upon by security professionals. Our collective discussion piece engages with the notion of “sensing security” (Klimburg-Witjes et al., 2021). In their book “Sensing In/security,” Klimburg-Witjes and colleagues foreground the socio-technical infrastructures that make specific securitizations possible. Our collective discussion advances this research agenda by examining the mediated practices of rendering security perceptible and actionable, and examining some of its political implications. More specifically, we contribute conceptually and empirically by focusing on three interrelated aspects: first, we examine how sensing technologies are understood to imitate or extend human perception; second, we consider how technological mediation simultaneously reduces and amplifies sensing and sense-making; and third, we analyze how these processes are deeply embedded in, and constitutive of, relations of power. After discussing some of the debates on sensing in security that this collective discussion is in conversation with, the introduction then proceeds to detail three core contributions.
Engaging with Literature on Sensing Security
In this collective discussion piece, the notion of “sensing security” provides a framework for re-orienting our analyses to explore the mediated ways in which insecurities are made perceptible, whilst also illuminating how security practices are enabled and constrained by them. One of the most important insights to emerge from CSS is how the perception and normalization of security threats takes shape in different practices (Bigo, 2002; Huysmans, 2006). These practices of securitization are characterized by deployment of an increased technological sophistication—ranging from surveillance systems, biometric scanners, to algorithmic prediction tools—that not only facilitate but significantly shape how security actors sense and make sense of security threats (Amicelle et al., 2015; Bellanova et al., 2021; Hoijtink and Leese, 2019a; Ruckenstein, 2023). In this collective discussion piece, we feel indebted to these approaches, including the focus on the materiality and agency of objects of security (Salter, 2015; Walters, 2002). Building on this scholarship, we, therefore, frame sensing security not as a purely mechanical or calculative practice, but as “a delicate interplay between humans, artefacts, and discourses” (Klimburg-Witjes et al., 2021: 28). More specifically, we argue that the value of “sensing security” lies in directing analytical attention toward the technologies and techniques that mediate both the perceptions and actions of security practitioners (Gabrys, 2016; Verbeek, 2005: 111–117).
In this collective discussion piece, we therefore analyze first how sensory technologies play a role in shaping how security is perceived. The existing literature in IPS and CSS already offers valuable insights. For instance, research on biometric border control demonstrates the work of detecting and verifying the identities of mobile populations (Glouftsios and Casaglia, 2023). Biometric technologies are designed to capture and display physical characteristics or parts of the human body that are not perceptible to the naked eye (Van Der Ploeg, 1999). However, how these sensory technologies depict their sensory input shapes how the body is perceived, interpreted, and acted upon and how they perceive border crossers (Valkenburg and Van Der Ploeg, 2015).
In addition, our contributions also shed light on the manner in which these technological mediators enable and constrain the actions of security practitioners. They not only impact how security actors engage with sensory technological mediations when perceiving and interpreting risks and threats, but also how they render these uncertainties actionable and governable (Amoore and Raley, 2017). This facet has been shown in the literature to be of particular pertinence when considering militarized sensing technologies, wherein strong ideals exist for seamlessly integrating the processes of perception and annihilation (Bousquet, 2018; Richardson, 2024). For instance, Gregory (2011) shows not only how remote drone operators engage with the video feeds from the aerial platform when perceiving and interpreting the battleground, but also what are influences of these “new visibilities” on military actions. He thereby argues that this sensing practice “produce[s] a special kind of intimacy that consistently privileges the view of the hunter-killer” with “implications [that] are far more deadly” (Gregory, 2011: 193). These sensory technologies thus shape not only what one perceives as suspicious, risky or a threat, but also who is the subject of surveillance and targeting. Although drone warfare might be experienced as a neutral intermediary by those using it—a more “precise,” “surgical,” or “objective” form of war conduct—it carries a form of sovereign power to decide the exception (Shah, 2012; Wilcox, 2017).
Engaging with the existing literature on the subject, the starting point for our understanding of “sensing security” is therefore that the knowledge about insecurities is not derived from direct perception; rather, what is experienced is always a mediated practice. Already, the “linguistic turn” in security studies taught us that there is no direct experience of something without the mediation through various discourses regarding security. For instance, the way in which border control, terrorism, drugs, organized crime, migration, asylum or the environment are perceived and treated politically depends on how they are represented discursively as dangerous, threatening, alarming and so on (Buzan et al., 1998). Building on these insights, more recent scholarship has focused on how human senses mediate perception (see also Bergson, 2012). Notably, seeing is often singled out, not only for being bound up with (Western) conceptions of truth and knowledge (Brighenti, 2007), but also a practice of surveillance (Amoore, 2009; Tazzioli and Walters, 2016). In addition, smell has been shown to be a form of embodied sensing capable of apprehending potential threat and enmity (McSorley, 2020). Moreover, Weitzel (2018: 421) has looked at “the central roles that sound, hearing, and voice” play in security practices. What we take from these different accounts is that there is not one dominant mode of sensing and perceiving a world that is to be secured, but there are multiple ways of sensing security.
However, as noted above, scholarship in IPS and CSS has increasingly de-centered the human subject and shifted its attention towards how technologies mediate sense-making. Perception is not limited to humans—nor to living beings, for that matter (Bourne et al., 2015; Rothe, 2024). In fact, the “sensory work” described in this collective discussion piece can “never be wholly distinguished from the material practices and modes of technological interface and analysis … via which they are produced” (Johns, 2017: 61). In the same vein, Suchman in the introduction to Klimburg-Witjes et al. (2021) calls for a focus on the socio-technical infrastructures of sensing “deployed in the name of securitisation.” Their edited volume shows the plethora of sensing technologies deployed for the purpose of security, surveillance and social sorting and, thereby, it also broadens the scope of security studies, where human language and imagery have long been focal points of analysis (Bleiker, 2018; Buzan et al., 1998; Callahan, 2020). A prime point of departure for our joint discussion is, therefore, to orient the analysis towards “sensing technologies” (Gabrys and Pritchard, 2018). Sensing technologies are not mere intermediaries, giving practitioners direct access to a world that needs to be secured. Sensing is a technologically mediated and relational practice, in which the experience of security becomes distributed among human and nonhuman actors.
Three Contributions
Drawing on this understanding of “sensing security,” our contributions in this collective discussion piece focus on three interrelated aspects. First, we seek to draw attention to how sensing technologies are often believed to imitate and extend human senses, despite the wider range of affordances offered by such technologies. For instance, surveillance cameras and dialect recognition systems are often designed to mimic human senses such as seeing and hearing, closely relating it to what is assumed to be a human way of making sense of the world. At the same time, despite being built with human senses in mind, sensing technologies introduce novel affordances. Many emerging sensing technologies go beyond what human senses can do. Technologies such as sonar, radar, infrared cameras, or (unsupervised) machine learning algorithms register signals that are undetectable by the human sensorium (Fish, 2022; Garrett and McCosker, 2017). The unaided and unenhanced human eye is ill-equipped to see “the forest for the trees” or to identify “the needle in the haystack” in large and complex data sets (Aradau, 2015). Sensor technologies often generate an excessive volume of data that is difficult for humans to analyze without leveraging data analysis tools and software. Therefore, despite the ideal to imitate and extend the human sensorium, the affordances of sensing technologies are qualitatively different from those of human senses.
The second aspect that is highlighted in our individual contributions on “sensing security” is how sensing technologies amplify specific aspects of reality while reducing other aspects of the world that need to be secured (Ihde, 1990). What these sensors detect as a security threat—and what not—is entirely dependent on the specific type of sensor solution that is applied to address the security problem at hand (Klimburg-Witjes et al., 2021: 24). In fact, each of the sensory technologies we discuss in our individual contributions prioritizes certain information (visual, tactile, and auditory) whilst ignoring others. Camera sensors, for instance, capture movements and transform these into machine-readable data. These inputs can subsequently be transformed into sensorial outputs (discursive, visual, and numerical) that security professionals can interact with. The point here is not that “real life” cannot be translated into technological format; it is that this translation can be done in multiple ways, and often requires multiple processing steps, which include and exclude certain aspects, with variable political effects (Scheel, 2019). Sensing, thereby, is productive of multiple worlds to be secured (Gabrys, 2016; Mol, 2002). In other words, our collective discussion underscores that the ways in which security is made sensible and perceptible is never a neutral process.
This brings us to the last aspect that connects the individual contributions, which is that sensing technologies also perform distinctive forms of power. That is, the power to discriminate against persons; to transform the way they are acted upon; shape the character of borders and warfare; and “(re)draw boundaries and lines of exclusion” (Amicelle et al., 2015; Rothe, 2017). Isin and Ruppert’s (2020) notion of “sensory power” captures this dynamic, whereby data practices chart a new and emergent form of power (see also Rouvroy, 2011; Amoore, 2020). With the increasing number of sensors that can produce data, and the infrastructures that can process this with new techniques of aggregation and correlation, our collective discussion piece underscores the argument that the “very experience of characterizing the world’s conditions, and of exercising power to govern, make legally significant decisions, … are under revision” (Johns, 2017: 60).
While our individual contributions do not go as far to make paradigmatic claims about shifting power relations, the collection attests to the fact that sensors that produce data for these algorithmic forms of government are emerging in many security-related contexts: ranging from battlefield, borders, to our very own universities. We build on the argument that algorithmic forms of security pose a serious challenge, for that they “necessarily involve new ways of reading signals, detecting disturbances and shaping action” (Amoore and Raley, 2017: 5). Indeed, the era of computers as passive data repositories and processors is behind us. Machine learning algorithms can interpret a variety of data and are increasingly able to act autonomously, “operat[ing] both beneath and beyond the threshold of human perception” (Amoore, 2020: 30; Aradau and Blanke, 2022: 172). However, we underscore in this collective discussion that sensing security does not have to be digital or automated; paper-based or “low-tech” forms of mediated sensing remain relevant in security settings (Bonelli and Ragazzi, 2014). Furthermore, we do not want to exaggerate the impact of the algorithmic sensing technologies in the fields we discuss by acknowledging their situated and contingent character (see also Aradau, 2023).
Our Collective Discussion Piece
To sum up, the goal of our collective discussion is to draw out some of these specificities of sensing security, which consist not only in showing how these sensing technologies partake in a re-ordering of perception, but also in exploring how they are productive of new practices of decision-making and securitization. The idea for this group discussion came from a workshop led by Klaudia Klonowska and Sofie van der Maarel in September 2023 at the European International Studies Association (EISA) conference. Afterwards, Jasper van der Kist, Natalie Welfens, and Ruben van de Ven joined. Through regular meetings and discussions about our respective contributions over a time span of more than one year, we developed the shared understanding of sensing security presented here, grounded in our diverse disciplinary backgrounds and empirical engagements. What unites us is a common concern with the non-neutral ways security is made perceptible and actionable through various forms of mediation. The concept of “sensing security” has proven to be fruitful in our collaborative discussion and individual inquiries—not as a rigid framework but as an invitation to focus on the ways in which perception, action and power intersect. This collective discussion is intended to push the debate on “sensing security” and to stimulate further thinking about the relation between sensing and securitization, as well as more research on mediated practices and technologies of sensing in security contexts.
The piece proceeds as follows. First, Van der Maarel discusses the military’s development of robotic technologies, emphasizing the focus on replicating human sensory capabilities rather than surpassing them. Second, Van de Ven examines the operations of “transduction” that allow sensor data of classroom occupancy to be computed and made actionable for the purpose of epidemic control in higher education. Third, Klonowska discusses “single pane of glass” as a primary interface for the military that allows for increasing amounts of heterogeneous sensors and data points to be algorithmically processed and visualized in a central location to identify unseen patterns and provide an operational view of the battlefield. Fourth, Welfens reorients the collective discussion towards embodied experiences of security, situating perceptions and judgments in refugee admission programs within the complex interplay of frontline workers and their mundane technological tools. Fifth, Van der Kist discusses the mediation of a particular sensing technology—speech biometry—in refugee status determination processes, emphasizing the discrete manner by which linguistic identities are made perceptible to decision-makers and how this has direct implications for how asylum seekers can present themselves discursively as a political-legal subject.
Envisioning Sensors as Extended Human Senses
Sofie van der Maarel
As an anthropologist studying security innovation through the framework of STS and CSS, I am interested in lived experiences of security practitioners working with new technologies. This section zooms in on military experiences of envisioning, developing, and testing sensor technologies. Specifically, it discusses how these technologies are envisioned as an imitation and extension of human senses to achieve an ideal of gathering and analyzing as much data as possible. Focusing on illustrations from a military case study, this contribution explores sensing in relation to military practices and perceptions of threat and enemy targeting.
Central in the desire of militaries to implement such technologies is its assumed potential to increase “situational awareness,” understood as the reading and making manageable of uncertain situations (Hentschel et al., 2025). Approaching situational awareness as a “governmental and everyday practice of decision-making and manoeuvring in highly uncertain settings” (Hentschel et al., 2025: 1), it also impacts the lived experiences of those working with sensing technologies. Reframing threat perception and response, these technologies shape assumptions of normalcy and security. As also discussed in the introduction, sensing is thus not a neutral or objective assessment of threat and risk, but a situated and embodied practice. This contribution further highlights the situated ideals of replicating human sensory capabilities.
In relation to the other contributions, sensing in this section revolves around gathering data before it is processed for decision-making practices. It draws on illustrations from a case study with the Dutch Army’s Robots and Autonomous Systems Unit. Soldiers in this unit experimented with semi-autonomous vehicles produced by Estonian company Milrem Robotics. This company focuses on “the synchronized deployment of soldiers, manned and unmanned air and ground vehicles, robotics, and sensors to achieve enhanced situational awareness, increased lethality, and improved survivability” (Milrem Robotics, 2023). According to the manufacturer, the vehicle is built for tactical reconnaissance missions and includes a “variety of sensors for day and night operations, acoustic gunshot detector, smoke screen protection and a ground surveillance radar [which] allows units to do multi sensor identification” (Milrem Robotics, 2023). The vehicle additionally contains two sensors to detect weather conditions: the wind direction and temperature.
Imitating Human Senses
At this Army unit, soldiers and manufacturers closely collaborated at a military base in the Netherlands to make the vehicle, which was earlier used in forestry, fit for military purposes. These actors decided that the vehicle must have two cameras that together would create 3D images. This idea was based on the nature of human eyes, in which the registrations of two eyes are made into one image in the brain. However, the vehicle does not have a human brain and the software behind it struggled to integrate two images into one coherent livestream. For people working with the vehicle, this was very confusing, and a lot of extra work was required to monitor, process, and integrate images from two screens. Such technological developments with the vehicle’s sensors followed the rationale of human senses. Whereas reasoning from the technological functioning, manufacturers mentioned that it made much more sense to have one fisheye lens camera, one “super eye.”
The extension of human senses through technologies was also discussed in relation to military maneuvers with the vehicle. As one soldier described:
Human senses have to be transferred to such a device. As humans we are pretty well developed, we register a lot. So far it has been difficult to get all our human senses in a device. For example, the vehicle was driving without a microphone, it was basically deaf. And it was being shot at from the flank and no one noticed it … If you are a human being walking there, you notice it if suddenly there is a sound coming from the side. With the robot … it does not see things like that … so one of the points we raised was to integrate acoustic sensors … seeing and hearing are the most important senses in technologies.
After that incident, the vehicle was equipped with four acoustic sensors that could pick up sounds and estimate what kind of weapon was used by potential enemies and at what distance. These sensors were not yet used in practice, but it remained soldiers’ wish to implement them in their operations to detect as well as to produce sounds. In that way, according to these soldiers, they can also be used to disturb the enemy and to be able to drive the vehicle towards enemies to pass messages, such as “surrender yourself.” The soldier explaining this admitted that this had “high killer robot” characteristics.
One military operator, moreover, stressed that the systems should add value to their operations in the form of extra eyes and ears, but “you shouldn’t be too busy with the thing.” If it becomes too complicated to operate, he stated, you can no longer look around and “you lose your own perception, because with your own eyes, you best see the enemy coming.” So, the sensors were only seen as valuable if they extended human senses, but were considered a burden when they distracted or diminished those. Operators thus ultimately wanted to maintain a sense of control by using their own senses to gain situational awareness.
These illustrations, in particular the fisheye camera and sound sensor examples, show how soldiers and manufacturers idealized sensing technologies as an extension of human senses, for example, more ears and eyes on the ground. Not the capabilities of the technologies were taken as a starting point, rather the focus was on how these technologies could mimic and extend humans’ physiological ears and eyes. Furthermore, many technologies were developed with the promise of distancing soldiers from the battlefield, offering the potential for remote operations that would enhance their safety (Van Der Maarel et al., 2025). However, the emphasis on the human senses in sensing technologies reveals a desire to keep the human involved, to gather data “as if” they were fully present. Almost as though soldiers were sensing with their own senses, only processing higher quantities of data. Realizing that this is unattainable as the large amounts of data of the “most complete image possible” cannot be analyzed by humans only, considering their cognitive limits, the focus shifted towards automation.
“Most Complete Image Possible”
While the focus in the unit was on sensing and the use of sensors in data gathering, or in military language, “tactical reconnaissance” or “situational awareness,” soldiers and manufacturers also discussed the analysis of this data. In doing so, they highlighted the importance of combining sensors. One soldier stated: “why do human eyes need to see it? Why don’t you automate that so that with an optical sensor you have an image, and that image is immediately recognized by the software.” He used an example of Russian troops:
Look, what you ultimately want is to generate the most complete image possible with such a [combination] sensor box. Now a soldier is sitting somewhere, and he observes, and hears or sees something and then the message is: “there are a thousand Russians coming towards me.” But the soldier 50 meters away says: “there are only 50.” So, what we want is that the data generated by the sensor box is as objective as possible and that it is therefore much more useful in the rest of the process. So, then you want to combine data, from the acoustic sensor, from the optical sensor, all kinds of things, and that all has to be integrated into a computer program. Well, where we’re at now is that all that data appears separately on a separate screen, and that is then analysed by some people who are knowledgeable, and they can make a complete picture of that.
These visions can be seen as ideals of sensing for security containing expectations, hopes, and desires to generate the “most complete image possible,” which is considered similar to the “most objective” image (this also links to Klonowska’s contribution on making security visible). Good sensing is thereby equated with more technologies, more data processing, and more integrated sensors. A combination of sensors and an integrated, automated analysis is considered more reliable than human judgment. Ironically, as mentioned earlier, human beings are the masters of “integrating senses” as our physical bodies are “programmed” to combine what we see, hear, smell, or touch to decide our course of action. So here again, the sensing technologies follow the logic of human sensing. According to the soldier mentioned above, the judgment would “normally” be based on two soldiers’ observations. The emphasis on combined sensors and automated data analysis suggests that this accounts for more accurate and “objective” sensing than those two soldiers (and all their senses) combined. These ideal visions of sensing closely relate to Haraway’s (Haraway, 1988a) discussions on the “god’s eye trick,” namely the desire and attempt to be able to see everything from nowhere. In this instance, it refers to the gathering and interpreting of information, whereas Welfens later discusses its relations to decision-making and positionality.
Moreover, focusing on how soldiers envision and design sensing technologies for security shows how these technologies and human-machine interactions are approached as an instrumental way to reach an enclosed world. This idea relates to the rationale that Suchman (2023) describes as ‘a resilient fantasy of data-driven, comprehensive command and control.’ Sensors, signal processing, and data transmission create an excess of data that threatens to destabilize the “technopolitical imaginary of just-in-time information, [and] AI is advanced as the promissory solution to automating data analysis and reclosing the world” (Suchman, 2023: 762). Soldiers’ discourses and ideal visions of sensing also evolve around gathering as much data as possible to gain situational awareness, or real-time understanding of what happens on the ground. Above that, a lot of trust is placed in automation processes to enhance soldiers’ sensory capabilities. Instead of altering how they see, hear, or sense, which is possible with novel technologies, these processes aim to augment their natural senses by gathering and analyzing more data rather than different types of data. From this perspective, sensor technologies are thus approached as imitations and extensions of human senses, and sensing security links to ideals of achieving security through generating the most complete sensing possible.
Ideals of Military Sensing
Concluding, this contribution explored sensing in relation to military practices and experiences of envisioning and testing robotic technologies. These were based on situated discourses and ideals, embodying military perceptions of what is considered a security threat. Compared to the other contributions, this section emphasized the human ideals of sensing technologies, showing how these humans approached them during their development in an instrumental and functional way. In sum, this contribution makes two observations about how security practitioners envision, design, and develop sensing technologies.
First, the human-nonhuman interactions in this illustration concern the idealization that robotic technologies should “sense” in the same way as humans sense. Following this logic, attention to sensing security in the military’s development of robotic technologies shows an emphasis on replicating human sensory capabilities rather than surpassing them. Second, this lens reveals the prevailing idea that the more senses are combined, the more objective the understanding of situational awareness. Hence, security knowledge is perceived as objective when mediated through combined and integrated sensing technologies.
These observations, however, reveal a paradox. As became visible, the more data is integrated, the more there is a need to use AI-enabled systems to analyze the large amounts of data. So, whereas the emphasis in envisioning and developing military sensors is on imitating and replicating human senses, this cannot be maintained for the analysis of the raw data. To turn this data into actionable information, military units such as the one discussed in this contribution rely increasingly on data analysis tools and software. As such, even though sensing technologies are often understood as imitating human sensorium, their purpose is to move beyond the human perception and register signals or patterns that are undetectable by the human sensorium (Aradau, 2015; Fish, 2022). Therefore, sensory affordances may be different from human sensing, which came to expression in the example of the fisheye camera. Envisioning sensors as extended human senses, thus, has its limitations for sensing security.
Sensing as Cascading Transductions
Ruben van de Ven
… if our ears were ten times more sensitive, we would hear matter roar– and presumably nothing else (Kittler, 2018: 352–353).
In her contribution, Van der Maarel considered ideals of sensing security for military data-gathering practices. In my contribution, I want to dig a bit deeper into this relation between data-gathering and sensing. For what are the politics at play when data-capture and processing become a sensing practice? In other words, what is at stake when we call something a sensor? By examining the constitutional moment of the sensor, I bring two things to the discussion. First, I want to engage in a relational account of sensing security that moves beyond the primacy of human-machine relations but also engages in the tensions that appear between various stages of machine processing. This will help me to upset the dominant narrative that sensors would provide “the most complete image possible,” as Van der Maarel’s respondent suggests, and instead tease out the primary function of sensors: a transduction of the environment, which is as much a reduction as an amplification of reality (see Mackenzie, 2002; Simondon, 2011). Like rose-colored glasses, sensors filter out much detail (which is considered noise), while painting everything in their color. Thus, examining the construction of a signal by tracing its transductions has a consequence for the dominant idea of the dashboard as providing an all-encompassing overview. This resonates with what Latour and Hermant (1998: 32) have called oligopticons, in which only “very little” is visible, but this narrow vision comes with “great precision.” The oligopticon is thus rather different from Bentham’s panopticon in which everything would be visible to the operator. As such, Latour proposes that both the megalomania of the operator and the paranoia of the watched are obsolete (lazy even!) responses to sensing practices. However, by examining the institution of such an oligopticon, I want to reconsider this implied depoliticization of sensing practices. Instead, I locate the sensor’s politics precisely in the negotiation of “very little” and “great precision” it enables; in the question of what is produced, what is filtered out, what remains in the processing of signals.
For this exploration, I turn to a case that happened at my own university in Leiden, the Netherlands. In November 2021, the student newspaper of the university, Mare, revealed that over the preceding months, the university had equipped many of its classrooms with devices intended to measure room occupancy in light of COVID-19 restrictions. Over 370 such devices—the PC2S developed by Xovishad been installed (Kloosterman and Reid, 2021). The university consistently called the devices “classroom scanners” and “occupancy sensors.” These devices, so Mare revealed in their publication, were in fact small cameras connected to a computer. They were able to quantify much more than room occupancy: built in algorithms to assess gender, race, and emotion of those passing by. After the publication by Mare, both students and staff rose up against these devices. Opponents critiqued the devices, calling them out for being surveillance cameras that pose a disproportional infringement of the privacy of any user of the buildings. The university board, in turn, apologized, not for installing the devices, but for their limited communication, implying the critique of these devices was due to a misunderstanding. After all, the devices should not have been considered as part of a surveillance apparatus, but merely as “sensors” that aid building management. Opponents, in turn, postulated that painting the cameras + computers as sensors was a rhetorical trick that downplays the surveillance potential of these devices.
What if, however, we take the framing of
camera + computer as a sensor herein seriously rather than
nefarious? This is not to simply sidestep the critique of these
“sensors” as constituting a surveillance infrastructure. On the
contrary, I want to use this case as a means to explore the ontological
politics that are at play when an ensemble of technologies becomes
designated as a single device. This is all the more relevant as this
case of referring to camera + computer as “sensor” does not
stand on itself. In fact, various manufacturers and professionals in
algorithmic surveillance do so5. The critique of the
camera + computer = sensor suggests that we can pry open
these sensors to further examine these security devices (Amicelle et al., 2015). Which relations
emerge if we draw out an exploded-view of the occupancy
sensor?
To understand how camera + computer makes a sensor, let
us briefly unpack the technological definition of sensors. One way to
describe a sensor is to consider it as a transducer: a device
that receives signals from a physical system, and which converts those
into other signals that contain information about that system. For
example, a temperature sensor changes its electronic conductivity
depending on the surrounding heat. The vibrating movement of particles
in the air is transduced into a variable electric current. The
technological feat of this transduction of one signal into the other
(the mundaneness of which only speaks to its marvel), thus, involves an
epistemological step of erasure: those interested in the temperature
signal are generally no longer concerned with the individual particle
movements that help produce the signal. Of concern is that which emerges
out of quantifying the kinetic energy of many particles: the
temperature of a substance is an emergent property. Likewise,
discussing the water control systems in Paris, Latour and Hermant write:
“Their wisdom is proportional to their deliberate blindness. They gain
in coordination capacities only because they agree to lose first water
and then most of the information” (1998:
28). A “catch-all” of data collection largely contains noise,
which needs to be discarded if one wants to be able to act on the data.
This seemingly resonates with the claim by the board of Leiden
University that the camera + computer = sensors do not
surveil individuals but “measure” the occupancy of rooms. Much like in
the example of the temperature sensor, room occupancy is merely
considered as an emergent property that follows from the movement of
people in and out of a space. However, following the work of Simondon,
media scholar Mackenzie (2002) argues that an attention to signal
transductions “aids in tracking processes that come into being at the
intersection of diverse realities. These diverse realities include
corporeal, geographical, economic, conceptual, biopolitical,
geopolitical and affective dimensions” (2002: 18). Transduction is not simply
reduction: at each step, the sensor filters, mobilizing assumptions
about the use of the data in order to throw away most of it, recasting
what remains in its own light. An attention to transductions helps to
outline the articulations of security that are at play in the production
of sensor signals.
Producing an occupancy “measurement” involves a huge number of
computer instructions, which I here bracket to five consecutive
operations. First, light hits a camera sensor, creating an electrical
current, which can be represented as a stream of bits to be
computationally interpreted: the signal produced by the image sensor.
One should not underestimate the quantity of data produced by an image
sensor: even a rather standard HD frame of 1920 by 1080 pixels adds up
to more than two million pixels. Assuming three colors, and eight bits
per color, this adds up to 6.22 megabytes per frame. At 30 frames per
second, such a device might produce over 180 megabytes per second.
Collectively, 370 of such image sensors produce over 66 gigabytes of
data every second. The electronics processing these signals simply would
not be able to keep up, were it not for the fact that image and video
encoders throw away vast quantities of the information. To do so, in a
second step, the JPG compressor (used in the common Motion-JPG video
format) transduces images into a series of wave functions, in a grid of
8 by 8 pixels. This removes fine details from the image, while keeping
the overall structure perceptibly similar to the input
image—“perceptually” here denoting the limitations of the human eye
(another, often used pixel format is YUV 4:2:0, which also employs a
reduction of pixels based on assumptions about the sensitivity of human
vision). In a third step, the camera + computer = sensor
takes this algorithmically processed image signal and predicts where
people occur in the image using computer vision techniques such as
convolutional neural networks. This creates another signal, one of
distinct detections. In a fourth step, the detections in one video frame
are compared with those of the previous frame using an object
tracker. In the fifth and final step, a fluctuating occupancy
signal can be produced by performing a rather simple integration of the
tracks into a numeric state: in and out of the room, +1 and -1. What is
left is the count, a numeric value, fluctuating in time: room occupancy
as a signal6.
Despite each transduction further obscuring the relation between the
individual subject and those operating the system (Aradau, 2023: 237), it is not
eradicated. On the contrary, as much as each transduction erases, it
also articulates this relation. Video encoders produce images for human
consumption, while the images are processed by algorithms, and
detections in individual frames are glued together based on assumptions
of how humans generally move. Each step of the process produces an
additional rendition of the individual. While all visitors to the
buildings were being filmed and counted by the
camera + computer = sensors, not all were equally governed
by this control system. The cascading transductions displace the
device’s primary subject of governance. For instance, on some occasions,
teaching staff got warning SMS messages indicating their room had too
many occupants, and they were likely violating COVID-19 restrictions.
University staff used to make their own judgment about the number of
students they let into their classroom; now, the agency of intervention
has moved to an automated system that signals the staff about occupancy
violations. This was not received well by many staff members, who
considered these interventions as an expression of sheer lack of respect
from the university board towards its staff. So, while the device’s
object detector is oblivious to classroom hierarchies—scanning students
and staff alike—(Lianos,
2003), the sensory assemblage imposes new hierarchies between
the operator and the subject.
In the case of the people counter, by placing “real time” sensors in classrooms, a logic of “smartness” enters the university building. As Halpern and Mitchell (2022) put it: smartness is an epistemology, a “knowing and representing the world so that one can act in and upon that world” (2022 xi). It was the continuous capture of sensor data that greatly increased the temporal resolution of the data available. With minute-precision measurements of occupancy, it became possible to not only optimize and anticipate, but also to intervene. The transduced signals form what Isin and Ruppert (2020) call “live clusters” that act as “intermediary objects of government between bodies and populations” through which such sensory assemblages facilitate real-time control. Sensing security is not just about producing highly specific measurements: constant, real-time monitoring often becomes convolved with (automated) intervention.
The relation between data-gathering and intervention lies not only in the immediacy with which it happens, but also in the precision that the sensor’s cascade of transductions affords. To make accurate control with the occupancy sensor possible, each individual needs to be detected, marked, and added to the occupancy count. To exemplify this, the story of Leiden University can be contrasted with another occupancy counter, one used by the primary Dutch railway operator. In their app and on their website, Dutch Railways indicates the occupancy of trains. They are able to provide this information without counting individuals entering the carriages. Already for a long time, trains have been weighed on particular parts of the railway track. The Dutch Railways use these measurements to estimate the number of people in a train. Such estimates suffice to monitor train usage but are not precise enough for direct intervention over the individuals on board. The university’s occupancy counters had been introduced to curb the COVID-19 crisis, thus their resulting signal—room occupancy—can be considered as a proxy for a property more relevant in light of the epidemic, namely the room’s air quality. Had the air quality been measured directly (such sensors do exist), an estimate of occupancy would have been possible. Nevertheless, such an estimate would not have been precise enough for direct intervention towards the teaching staff. In short, the accuracy with which signals are produced is of relevance to the governing practices afforded by sensing security.
I have discussed sensing security as an attention to the
transductions that are happening in the various stages of signal
production. While “sensors” are generally conceived as mere input points
of data, they are devices that take on computational work (Pelizza and van Rossum, 2021). This
attention to the sensing device counters the depoliticization of mere
“measurement instruments” in two ways. First, much like Latour’s take on
the oligopticon, tracing transductions renders the assemblage as an
exploded view, allowing us to inspect the relations that are drawn into
the reduction and amplification of reality. This is what we saw in the
initial critique of the people counter: calling the device a sensor
brushes over the complex sequence of operations that takes place. Thus,
opening up the occupation sensor forces acknowledgment of its
composition as camera + computer; the sensor is a device
that produces a video stream. The 370 sensors are 370 cameras that
become latently part of an already existing infrastructure for
surveillance. Thus, the composition of the sensor device matters, for a
politics of the sensor does not replace that of its components, it
multiplies them. Second, tracing the cascading transductions of the
camera + computer = sensor locates the politics of its
aggregated signal in the negotiation of immediacy and precision as
facilitators of intervention. After the uproar, the university board
decided to take down the camera + computer = sensors. To
still get a grasp on room occupancy, the faculty are back to sending
around people with a notebook. At set intervals, they briefly interrupt
presentations and conversations by opening the door and counting the
people in the room. Like the “occupancy sensors,” they produce a numeric
outcome. But, relying on their human forgetfulness, they are seen by
many as less of a privacy hurdle. They rely on sight, but not on the
production of image artifacts. Moreover, working in silence, the human
counters abide by traditional classroom hierarchies and are perceived as
less threatening to the teaching staff’s authority. The real-timeness of
the measurements has decreased, but usage of classrooms can still be
optimized and anticipated, while staff have regained the agency for
intervention.
’The Single Pane of Glass’: Making the Battlefield Visible and Actionable
Klaudia Klonowska
Other interventions in this collective discussion attend to the diversity of sensory inputs and the computational operations that transform sensory inputs, such as sound bites or electromagnetic waves, into readable signals. I suggest that the study of sensing security additionally entails attention to interfaces, or in other words, ways in which computerized sensory inputs are made visible and received by human users. Interfaces bring attention to how visibility becomes simulated and “titled phenomena appear coherent while obscuring the layered, socio-technical operations that produce and reproduce them” (Johns, 2017: 10–11). In this short intervention, I am primarily interested in how aspirations for greater situational awareness are operationalized, what picture is attained, and how it is perceived and made actionable.
I come into this collective discussion with a background in international humanitarian law and an STS-inspired interest in what security assemblages enable in terms of legally relevant military decisions and conduct. Though “sensing” is often overlooked in legal research, several authors have underlined its relevance to matters of justice and governance. Most notably, Johns (2017: 60) has argued that jurisdiction gets exercised only when the legally authorized agent “gain[s] both the capacity and inclination to sense a certain condition, person, or pattern.” Hamilton et al. (2017: 2) further argued that vice versa, available “[w]ays of sensing shape ways of knowing and the regulatory projects upon which they rely.” As the intervention of Van de Ven shows, attention to the socio-technical practices and complex transformations that take place when sensory technologies are employed provides opportunities for critical engagement with the implications to the exercise of human rights. In these spaces, lawful interventions of security actors rely on the capability to visualize threats. As this contribution will further argue, paying attention to the interfaces of sensing technologies and their users can tell us more about the distribution of authority and power to intervene.
In this piece, I draw insights from interviews with practitioners of the United States Task Force 59 (TF59) conducted between March and May 2023. TF59 is a naval unit with two operational hubs in Manama, Bahrain and Aqaba, Jordan that experiments with and deploys cameras, sonars, thermal radars, aerial and underwater drones, unmanned vessels, and algorithmic tools to sense the waters of the Red Sea, Gulf of Oman, Persian Gulf, Arabian Sea, and parts of the Indian Ocean (Brandi 2023). The main purpose of this extensive endeavor—which includes one of the world’s biggest fleets of unmanned vessels—is to detect and identify threats, and in turn improve maritime domain awareness. TF59 serves as an excellent illustration of a broader trend of how the selection of “relevant” data and the performative associating of data into profiles contribute to the coherence and operability of the common operational picture.
“Single Pane of Glass”
One of the major paradoxes of the digital era is that the volume of data that we produce and collect is so large that it is neither manageable nor digestible by humans alone, which also became clear in Van der Maarel’s case. While data-gathering techniques improve, new solutions such as machine-learning algorithms are needed to manage its volume and represent data in an accessible format. Vice versa, as data-hungry machine-learning systems are developed, the need for data-gathering grows in an endless cycle. For many security actors, a solution to the problem of data overload is the design of databases and visual interfaces that serve as access points. At TF59, this solution is called a “single pane of glass.”
A single pane of glass is based on the idea of a visual representation of an operational environment in which data from all kinds of sensors and data-gathering devices mounted on unmanned vessels are pulled together to create a uniform and coherent picture of what is happening “out there” on the open seas. This picture is based on signals from a wide variety of sensing devices, including satellites, drone cameras, and multibeam sonars. The term “single pane of glass” is meant to highlight that—instead of looking at the world through several different computers and “windows”—this technology enables the military to rely on one interface, one “window” through which visualization of the operational environment occurs. In this process of data centralization and visualization, the aspiration is to improve the common operational picture of security actors, meaning their shared awareness and understanding of emerging threats in the maritime domain.
The choice of the single pane of glass entails that the multi-model sensorial data that is gathered from various devices is reduced to a uniquely visual form of representation that relies on pictures, maps, colorful graphs, indicators, and heat maps. Sometimes access to sound may be retained, but likely the single pane of glass prioritizes the visual at the expense of other senses that are made remote. Though it should be of note that human operators can follow up on the recommendations presented by the single pane of glass and deploy a patrol boat to waters to gather “more information” using their physiological senses. The single pane of glass is not an accidental form of representation of senses; rather, it shows that the well-reported phenomena of “ocularcentrism” still live on, described as a tendency of humans to prioritize the visual experience in making evidentiary representations of the world (Hamilton et al., 2017: 4). These evidentiary representations, which visualize risks on the waters, then mediate the sensing experiences of human operators and inform their decision-making processes.
Constructing a Coherent Picture
The creation of the single pane of glass, however, is not a straightforward process of representation of what is “out there.” Instead, it is a complex process that depends on both human expertise and choices, as well as algorithmic processing and visualization practices that construct what appears to be a complete picture. In this process, selection of sensing devices, collection of data, and choices about risk profiles contribute to the sense of coherence of the single pane of glass.
The main challenge for the creation of the single pane of glass is the choice of which data is “relevant” so as to be visualized and presented to the human operators. Machine-learning algorithms are critical to this process. Amoore (2020: 9) remarks: “the very essence of algorithms is that they afford greater degrees of recognition and value to some features of a scene than they do to others.” Algorithmic filtering is thus a process of selection whereby certain data is rendered “relevant” and visualized. As opposed to filtering, which relies on the removal of unwanted information, selection highlights the process of adding together. What “relevant” in these contexts means is an outcome of socio-technical chains of interactions, influenced by the trade-offs of engineers in the algorithmic model design, abnormalities in the emerging patterns compared to “normality” observed in human pre-selected training data (see also Pasquinelli, 2019). Ultimately, social ideas about what a threat “looks like” in these security contexts will influence the design and iteratively selection of “relevant” features.
In the case of TF59, the selection process is aimed at automating the detection of maritime vessels exhibiting “abnormal” behavior, indicating potential smuggling of weapons or drugs, piracy activities, human trafficking, or illegal fishing. To identify and visualize “abnormal” behaviors onto a single pane of glass, the system swifts through data points from the Automatic Identification System (AIS), which is a database that has been used to monitor maritime traffic since the early 2000s. AIS data is a numerical geolocation output produced by the collection of satellite imagery, radio signals, representation of location on maps, and human input. The use of AIS for the analysis of maritime traffic has a long history, but the use of machine-learning algorithmic techniques to automate the processes of sensing and knowing which vessels to look further into is relatively novel. Through this automation, the promise is that more vessels with abnormal and suspicious activities are detected so as to enable lawful interventions. In this process, the algorithm creates added layers of complexity to identify “abnormal” patterns: it uses both human-selected “features” (such as the behavior of two vessels meeting at sea which for long has been considered to indicate risky behavior such as smuggling of materials), as well as self-generated “features” that the algorithm develops while “learning” from training data and subsequent use.
Counterintuitively, the creation of the single pane of glass does not only rely on existing data, but in order to provide a sense of coherence it also relies on the predictive capabilities of machine-learning systems. As algorithms selectively visualize relevant data, they also simultaneously constructively fill in the gaps. The creation of a seemingly coherent picture that informs how vessels are to be classified into risky/non-risky or normal/abnormal maritime behaviors relies on the associative power of machine-learning algorithms to produce profiles.
For example, data from satellite imagery is collected in up-to-15-minute intervals, which results in temporary invisibilities of the movements of vessels on the seas. Machine-learning algorithms are used to fill in those invisibilities with visibilities, in other words, to provide a visual representation of a possible direction of movement of those vessels in the periods where sensory signals are missing. This is done by comparing common patterns of movement (in specific routes, for a specific type of vessel, in specific regions) and predicting the possible pathways that the vessel might take. In this way, despite missing data, the vessel’s movement is estimated and used to determine its classification as normal/abnormal. By filling in the gaps in the movement of vessels, the machine-learning algorithm contributes to the drawing of trajectories onto maps and graphs so that objects can be visualized and sensed by human operators.
The result is a seemingly coherent and operationalizable picture that can be used in a decision-making process. The creation of a single pane of glass, however, is not an objective representation of reality; rather, it is a process through which new layers of visibility are produced. The co-production of visibilities is where “relevant” threats/risks/bodies are defined, and gaps and invisibilities are filled by the meticulous work of the algorithms, technical developers, and human users. In this process, the technology redistributes which vessels can be sensed by security actors, how they are made sense of, how many spaces can be sensed at the same time, and, therefore, how many bodies and movements can be governed and acted upon at the same time.
The Power of Differential Visibilities
The choice of creating a single pane of glass is meant to serve as TF59’s common interface for the detection of threats on the open waters. It gives the impression of centralization. Indeed, the technical pursuit of a single pane of glass means that data formats are standardized, and data is stored in shared clouds. The centralization of data through cloud computing also means that it can be easily accessed by different units, making the single pane of glass available to various security actors. The aspiration is that command posts, as well as different deployed naval units and other states or partner organizations with whom they are cooperating, can all have access to the same data infrastructure and act upon a shared (meaning standardized) understanding of the emerging threats. Naturally, this aspiration is not Navy-specific; similar projects of creating the single pane of glass are underway in the Army and Air Force. For example, the US Army is advancing its “deep-sensing” initiatives to improve the “ability to see and sense deeper, delivering an organic collection capability” (Baldwin, 2024; Klonowska and van der Maarel, 2025).
Sensing is often presented as an organic, natural, and even neutral act of information collection. This understanding, however, downplays the important connection between data accumulation and information processing capabilities on the one hand and the ability to “act on a distance” on the other (Miller and Rose, 1990; see also Van Der Kist et al., 2019). In his book Reassembling the Social, Latour (2005: 181) notes that the command post’s capacity to control stems from their ability to maintain such large networks of circulating references accounts for an essential form of power:
It is obvious, for instance, that an army’s command and control centre is not “bigger” and “wider” than the local front thousands of miles away where soldiers are risking their life, but it is clear nonetheless that such a war room can command and control anything—as the name indicates—only as long as it remains connected to the theatre of operation through a ceaseless transport of information.
In line with Latour’s argument, the informational-material arrangements that condition and make it possible to “see” and act upon the operational environment from one single command point, as long as the “connections” through which information travels “hold” (Latour, 2005: 181). From this point of view, my example of the single pane of glass illustrates how this sensing technology is intricately related to the accumulation of power.
Whereas the technological developments are pulling in the direction of greater diffusion of control centers, at the same time, the military institution, married to the ideas of hierarchization, implements new strategies to control the flows of data. One such strategy is the layering of the single pane of glass. This is also another means by which interfaces and sensor data are made operable. The single pane of glass is meant to be scalable, so that actors can be assigned different access depending on their clearance level. Such representations are then multi-layered so that a commander has access to more detailed information versus an operator whose access may be limited. A corporate partner may also be excluded from access to classified data and hence may have access to only a certain layer of the single pane of glass. This layering feature allows the military to manage which information is visible to actors that are connected to the system, and thereby also manage the control centers. Hence, attending to the interface and its various technical design choices brings attention not only to what information becomes sensible but also to how the capacity to act upon that information is arranged through the material-semiotic means.
To conclude, sensing security is not only a way to make objects/persons sensible but also a socio-technical means to distribute authority and power and make those objects/persons governable. Hence, attention to sensing security also entails attention to interfaces, which are the moments of encounters between computerized forms of sensed data and human users mediated by technologies, context, biases, language, and a variety of other subjectivities and materialities that make sensing and knowing situated. The turn to sensing also entails challenging technological determinism, a view that technology “as an inherent logic, insusceptible to external factors” by tracing human-nonhuman configurations and “reverse blackboxing” how things are rendered visible and invisible in those complex networks (Schiølin, 2023: 368), and how different interests are involved in the making of security threats. There are socio-politico-legal questions within each transformation. Scholars with a particular interest in legal considerations can attune to injustices that are produced through sensing regimes, when visibilities and invisibilities are stabilized (see Van der Kist’s contribution), when calls for transparency are mobilized, when resistance to visibility is expressed, and when access to sensing and hence knowing is unequal.
Sensing Security in Refugee Resettlement: Low-Tech Tools, Human Discretion, and the Politics of Categorization
Natalie Welfens
In my research, I am interested in categorization and sorting practices in the context of refugee and migration governance and the social inequalities they produce. In this field, just like in other security contexts, the use of technologies is proliferating (Ozkul, 2023), triggering questions on, for instance, how technologies construe migration as a threat (Huysmans, 2000) and policy problem (Glouftsios and Scheel, 2020), or reinforce epistemic violence in asylum procedures (Scheel, 2024). Looking at other security fields, the contributions to this collective discussion so far have illustrated how “sensing security” can offer a fruitful analytical lens to render visible how technologies mimic and extend human senses, the politics of sensors as transducing devices, and the processes of visualizing sensory data to provide operational knowledge. These empirical examples have underscored the role that high-tech sensors and their algorithmic logics play in mediating security and rendering it actionable, but also the role that human interpretation of data and signals play.
In my contribution, I would like to deepen and complicate these conversations by focusing on low-tech tools and discretionary sense-making by human agents in the context of refugee governance. First, while high-tech sensors often dominate discussions of sensing security, my analysis highlights how also mundane inscriptions (forms) and technologies (phone calls) mediate categorizations of refugees as ‘at risk’ or ‘a risk’, as objects of care and/or of control (Aradau, 2004; Pallister-Wilkins, 2015). These administrative tools, though seemingly simple, are central to reducing the complex world of mobility and security into one-dimensional categories in order for it to be made sense-able and governable (see also Bonelli and Ragazzi, 2014). As such, they impact whose protection needs are heard, seen, registered, and addressed, and which ones are ignored. Second, my contribution sheds light on embodied practices and highlights the situated knowledge of frontline workers as a critical aspect of “sensing security.” This resonates with the observation that while technologies in migration governance proliferate and produce “control knowledge,” human case workers still often assume the positions of “epistemic gatekeepers” (Amelung et al., 2024; Scheel, 2024). Interrogating these aspects of human discretion in sensing security, I argue, is key for locating power in human-technology configurations (Suchman, 2006), and when examining the “redistribution of the sensible” (Rancière et al., 1999), as Van der Kist will expand on in the subsequent contribution.
The empirical context in which I will illustrate this argument is refugee resettlement. Through resettlement and similar refugee admission programs, Global North states admit limited quotas of the “most vulnerable” refugees from countries of first refuge, while excluding those they deem “risky” for the nation state. In selection practices, categorizations serve to assess, on the one hand, risks to refugees’ human security (e.g., exposure to violence, lack of access to healthcare), and on the other hand, the risks they may potentially pose to admission states’ security (Welfens and Bonjour, 2021). As resettlement is a discretionary, humanitarian commitment by states, spots are scarce: less than one percent of the world’s refugee population will be resettled. Therefore, categorizations of refugees’ risks and riskiness not only powerfully illustrate how classification systems stratify, include, and exclude (see also Bowker and Star, 1999), but also their real-life consequences. For that, as (Fassin, 2007: 500) notes, “the saving of individuals, […] presupposes not only risking others but also making a selection of which existences it is possible or legitimate to save.”
Building on Latour’s (1999) concept of the “chain of translation” and its development in IPS (De Goede, 2018), I conceptualize refugee admission as a chain of admission: a transnational process through which individuals in countries of refuge are transformed into legally recognized refugees via a series of interpretive and bureaucratic steps . At each link—spanning NGOs, international organizations, and state agencies—categories like “vulnerability” or “risk” are not simply applied but translated. As dossiers circulate, categorizations are reinterpreted through situated judgments and shifting institutional logics (Lipsky, 2010; Yanow, 2003). This translation is mediated through both low-tech tools (e.g., paper forms and spreadsheets) and more advanced technologies (e.g., digital case management systems, databases). Sensing security in this context then encompasses interpreting and translating this technologically mediated information by caseworkers based on their tacit or situated knowledge—i.e., “the very mundane, but still expert, understanding of and practical reasoning about local conditions derived from lived experience” (Yanow, 2003: 236).
These dynamics become tangible in the everyday practices of refugee resettlement, where sensing, categorization, and technological mediation converge in concrete bureaucratic processes. To illustrate, based on fieldwork in Turkey in 2019, I will focus on one particular link in the chain of admission: when referrals of people “in need of resettlement” reach the United Nations High Commissioner for Refugees (UNHCR), case workers make an initial assessment of people’s resettlement need and eligibility by means of a phone interview. Unlike the high-tech sensors in military operations and pandemic control described above, UNHCR, in its assessments of referrals, relies on more low-tech technologies, which, however, also mediate sensing in particular ways. For instance, UNHCR Turkey receives the list of initial referrals—Syrian refugees identified by the Turkish migration authority as having a “resettlement need”—in the form of a basic Excel spreadsheet. Next to a few columns with personal data, the reason for referral is reduced to one category, such as “medical needs,” “women in need,” or “survivor of violence or torture.” In this way, initial vulnerability assessments and the complex accounts of refugees’ situation in Syria and Turkey are condensed to a single line in an Excel spreadsheet, filtering out what is considered irrelevant “noise.” In that way, refugees’ lived experience is reduced to bureaucratic categories, stripped of its complexity and nuance, while amplifying what, from UNHCR’s and states’ perspective, enables the bureaucratic governance of mobile populations.
Such processes of reduction and amplification are not neutral and have significant real-life consequences. For instance, when resettlement of Syrians to Europe started, Turkish migration authorities had relatively little experience with assessing and classifying cases, which led to high numbers of referrals under the category of “medical needs,” as “these are the vulnerabilities that are often the most visible, the ones that the eye can see,” Mahmoud, the UNHCR case worker explained to me. Here, visibility, methods of measurement and their translation into technologically mediated categories are crucially political: what is deemed visible (e.g., physical injuries ) is what gets counted, while other forms of vulnerability (e.g., psychological trauma and social marginalization) are rendered invisible. At the same time, cases classified primarily as “medical need” can be particularly difficult to resettle, as many admission countries are reluctant to take substantial numbers of refugees with medical conditions. Even though these cases are clearly “at risk,” they are also considered “a risk” for admission states’ economic and welfare security.
UNHCR case workers like Mahmoud call the “resettlement candidates” submitted to him in an Excel spreadsheet one by one. The initial phone interview serves to translate the broad categories from the Excel sheet into UNHCR’s own digital case management file and to classify the cases along some of the criteria that UNHCR deems most relevant for further assessments: the protection needs people describe, whether they want to be resettled, where extended family lives and what documents they possess. At this link in the admission chain, through the digital case management system and its categorization logics, the various information gets reduced to ticking boxes. For instance, in the administrative form in UNHCR’s digital case management system the question, whether the family wants to resettle only allows for a yes/no answer—although refugees may not know what resettlement is, cannot choose the admission country and may first need to discuss the option with their family.
Initial phone interviews also serve to classify cases’ urgency. Mahmoud explains how his intuition and tacit knowledge, shaped by his Syrian background and lived experience, help him “sense” urgency in a caller’s voice. This practice extends human auditory perception, transforming spoken words—for example, a refugee’s description of their situation in Turkey and immediate needs—into actionable data such as classifying a case as “medium urgency.” This, in turn, indicates to the next case worker in the chain how swiftly they need to do a follow-up interview. This interplay between fixed categories and human discretion underscores the tension between technological standardization and human agency. While the Excel sheet and digital systems produce a durable, authoritative framework for categorization, Mahmoud, as a frontline worker, acts as an epistemic gatekeeper. His situated knowledge—his Syrian background, his fluency in Arabic, and his familiarity with the cities refugees fled from—enables him to interpret the “noise” that the system filters out. For example, when Mahmoud speaks to a Syrian woman on the phone and classifies her as “at risk” with “medium” urgency, he puts her on speaker so that I can listen to their conversation, though without understanding the content. It is in this interaction with reductive inscriptions and devices that Mahmood’s perception and cognition are amplified, allowing him to govern and decide whose cases count as urgent or not, which in turn significantly shapes refugees’ chances to be resettled.
Epistemic gatekeeping and the interplay of human agency and low-tech tools also matter in interpreting standardized rules for prioritizing and “deprioritising” cases. In UNHCR’s training manuals, there are some fixed categories, almost following an algorithm’s “when A then B” logic. For instance, at the time of my observations, when women were “without male protection,” they were automatically classified as “women at risk.” In contrast, someone who served in the military after 2011 in Syria, served for the Ba’ath party or had a high position in the government was automatically flagged as ‘a risk’ and deprioritized for resettlement. Similarly, cases of child marriage and polygamy are, despite their vulnerability, flagged as deprioritized in UNHCR’s digital case management system, making these people’s resettlement highly unlikely. Also here, the complexity of a case gets reduced to a bureaucratic marker, signaling a person’s risk and riskiness. Yet, the fact that translating refugees’ accounts into bureaucratic categories is human—not machine-made, allows for situated interpretation, weighing of factors and, importantly, exceptions to the rules. As Mahmoud explains to me, “For all deprioritisation cases there is an exception, because life matters.” For instance, when there is an urgent medical case, UNHCR still tries to process the case.
It is precisely such elements of discretionary judgments and sense-making that risk to get lost in algorithmic and, however rare, fully automated decision-making. In Germany’s federal admission programs for vulnerable people from Afghanistan, for instance, access to the program was largely determined by an algorithm-based scoring system. Instead of assessing individual applications based on detailed, case-specific justifications, eligibility was pre-filtered through a digital questionnaire with yes/no responses and box-ticking answers. Civil society actors filled in these questionnaires on applicants’ behalf, without always having a sense of how factors would be weighted and how the scoring system operated. Reviewers then only examined cases pre-selected by the algorithm, grounding their assessments on brief summaries, without provisions for nuanced evaluations and, notably, exceptions to security-related exclusion.
To conclude, as highlighted in the introduction to this collective discussion, sensing security as an analytical lens invites to interrogate the power dynamics embedded in sensing practices, including the question of how technologies redistribute power and authority. As the example of resettlement practices illustrated, this focus also matters for examining low-tech tools and the way they shape security perception and actions. Like high-tech sensors, they are crucially political in that they produce hierarchies of visibility, reduce abundance of data to one-dimensional categories and thereby “make up people” (Hacking 2006)—here, as promising “resettlement candidates” or deprioritized “security threats.” Yet, these low-tech systems, at least at some links of the admission chain, attribute substantial epistemic gatekeeping to humans and situated interpretation, compared to some black-boxed AI-driven systems. As Mahmoud’s practices illustrate, interpretation and situated judgments are informed by positionality, familiarity with local contexts, the ability to “read” a caller’s voice and, where necessary, overturn “when A then B” decisions. This embodied knowledge contrasts with the rigid categories of semi-automated and algorithmic systems that substantially reconfigure how power is distributed between human and technological actors—a theme Van der Kist will further explore.
States of Perception: On the Redistribution of Sensing Security
Jasper van der Kist
In my research, I am interested in the entanglements of epistemic, technological and legal practices of refugee protection, particularly in relation to the procedures of refugee status determination. As noted by Welfens in her contribution, the fundamental principle underlying the operation of street-level bureaucracy is that it requires individuals to make decisions about other people (see also Lipsky, 2010: 161). While these discretionary practices have always been materially mediated in one way or another (Magalhães, 2016; Van Der Kist, 2024, 2025; Wissink, 2021), asylum procedures are becoming increasingly saturated by technologies affording new and divergent ways of sensing, perception and action (Scheel, 2024). In this contribution to the collective discussion piece, I will further underline the previous discussions on the mediated nature of sensing, in particular with regard to the generally overlooked auditory faculties (Ihde, 2007; Weitzel, 2018). Moreover, I would argue that a focus on sensing security has important political-legal ramifications, because it orders what can be said and perceived, and by whom (Rancière et al., 1999).
The State of Perception
Given the general tendency of officials to view the asylum seeker’s words with suspicion, and the perception that these words are unreliable and distorted, asylum administrations have long sought to outsource or develop in-house alternative forms of evidence (Lawrance and Ruffer, 2015). To assess the risk of persecution of the asylum claimant, for instance, medical techniques of sensing the body have long been used in asylum procedures to detect tangible signs of physical trauma (such as scar tissue) on the body of the migrant. Consequently, in this context, the concept of sensing security can be likened to the biblical narrative of Doubting Thomas, in which the sovereign is compelled to physically inspect the wounds of the stranger in order to validate their claims (Fassin 2012, 113).
In recent years, however, governments have introduced a plethora of computer-based and data-driven technologies ranging from physical and behavioral biometrics, automated document verification to the use of various predictive assessment and triage tools (see Ozkul, 2023 for an overview). This has mostly been motivated under the guise of creating a fairer, more efficient and effective asylum system; technological systems that can offer asylum administrations neutral solutions to a governmental problem, objectively perceiving the identity and risk of the migrant. Contrary to these views, however, I argue that the implementation of such technologies is never “innocent”—to invoke another one of Haraway (1988a) concepts. As technological mediators in their own right, these technologies introduce new and divergent affordances—that is, they enable and constrain alternate modes of sensing, perception and action (Gibson, 1979). More importantly, as I will demonstrate below, they may do so at the expense of other modes of experience.
Certainly, the contributors before me illustrate how sensing technologies are capable of perception and action. What has been implicit in the discussion, at least conceptually, is how sensing technologies relate to security decisions. The empirical examples given portray sensory prostheses to make strategic decisions on the “let live and make die” (see Van der Maarel); territorial “command-and-control” (see Klonowska); “classroom sensors” for university management to act on the unruly behavior of students and staff (see Van de Ven); or devices for administrators to determine whether risks should be attributed to the foreigner or the state (see Welfens). I would argue, therefore, that sensing security is intimately associated with sovereign practices of decision-making—or rather “any institution having the power to relate consequences to perceptions” (Stengers, 2018: 60).
What has remained implicit in the discussion, too, is how security perceptions are redistributed by the new “algorithmic” modes of sensing, and with what political consequences. This argument should not be confused with an all-encompassing “technological hype,” as we know from genealogical works, for instance, that the introduction of sensory inventions like the stethoscope gave rise to a distinct perception of sick bodies, changing the power dynamic in medical hospitals by subjecting patients to the authoritative gaze of the doctor (Rice, 2008). With the current invention of computational and data-driven technologies, we may be witnessing such a “redistribution of the sensible” (Rancière et al., 1999). What is important in a research agenda of sensing security is therefore not only what phenomena are made perceptible (and governable) by these new sensory technologies, but also remain sensitive to political questions of who (or what) remains imperceptible (see also Amoore, 2020; Aradau and Blanke, 2022).
I will limit my contribution here to a brief description of one of the technological developments in refugee status determination procedures in Europe, namely automated dialect recognition (Van Der Kist, 2025). In order to illustrate two points, I will present an example taken from a pilot project in the German context. First, how the design of this auditory technology orders the perceptions and actions of caseworkers in unequal relation with the asylum claimant. Secondly, how it redistributes who or what can (or cannot) be said or perceived in these administrative-legal procedures.
Making Linguistic Identities Perceptible and Actionable
In the aftermath of the so-called “migration crisis” in 2015, the German Federal Office for Migration and Refugees (BAMF) developed the Dialect Identification Assistant System (DIAS). This is a data-driven technology designed to assist authorities in processing asylum applications in a more efficient, secure and accurate manner (BAMF, 2018, 2022). At its core, DIAS employs speech analysis technology to assess whether the accents and dialects of asylum seekers align with their stated countries of origin. Similar to Van de Ven, I want to start by “unpacking” the automated dialect recognition device itself, unveiling how the identity of a migrant’s voice is rendered perceptible and actionable.
There are several intermediate steps between the registration and identification of the asylum seeker. The first step is to make the voice “machine readable” (van der Ploeg, 2005). By asking the asylum seeker to speak through a phone, a sensor scans their voice characteristics and generates a datafied version of the speech sample of it that is then stored in a local database from the asylum administration. The second step consists of the processing of this speech sample, where its features will be compared against templates stored in the database. The latter database consists of language samples of native speakers that had already been collected beforehand. Moreover, a model has been created based on machine learning, which is able to predict relations between them. The third and final step is the generation of an output, an electronic document with a comparison score measuring the degree of similarity between the speech sample and the template. For example, it may indicate that, according to the algorithmic assessment, there is a 54 percent confidence that the asylum seeker speaks Levantine Arabic.
It is through these many intermediate steps that the migrant’s voice—which is assumed to have unique biometric features—is being “made calculable” by caseworkers in the asylum administration. This enables a highly specific way of perceiving the asylum seeker’s identity that may have a significant impact on the outcome of their application. It is therefore crucial to acknowledge, as demonstrated by Van de Ven and Klonowska, that the process of identification entails a transformation—or “translation” (Callon et al., 1981)—whereby certain information is filtered and lost in order to gain these new insights (Bellanova and Fuster, 2019: 360). The final outcome of this process is contingent upon a multitude of factors that can influence the outcome and alter the final result (Callon and Law, 2005). For instance, asylum procedures are predicated on the highly contentious yet bureaucratically entrenched categorizations of identity, which in the case of automated dialect recognition are also those founded upon the assumptions of “dialect geography.” Embedded in the device is, therefore, the (methodological nationalist) presupposition that the language one speaks is tied to one’s place of origin. They are not sensitive to the wide variety and mutability of dialects. This assumption can become problematic in real asylum cases because, as Pfeifer (2023: 196) notes, “migrant biographies often involve spending months or years in transit, which means that people become immersed in different linguistic communities and contexts”.
Redistribution of the Sensible
The point here is not that automated dialect recognition system fails to represent asylum seeker identities. Rather, what interests me is that the sensing technology requires its own particular translations that pattern the perceptions and actions of decision-makers in ways that are somehow different from what came before. For instance, the perception of asylum seekers in conventional administrative-legal procedures is in large part determined by listening to the asylum story. Whilst it also depends on auditory faculties of the caseworker, this spoken, dialogic form of mediation diverges from those of automated dialect recognition (Van Der Kist, 2025). The automated dialect recognition systems are not interested in “what asylum speakers actually say” but are rather “grounded on the assumption that whatever they might say is not worth being taken into account before the trustworthiness of their belonging is ascertained” (Bellanova and Fuster, 2019: 360). Moreover, its calculation does not offer any understanding or explanation for the output. Rather, the algorithmic approach to speech simply suggests a correlation between the language model trained on an existing database and the speech data extracted from the asylum seeker. Other more complex and nuanced aspects of identity, including biographical markers, and how asylum seekers narrate them, are rendered irrelevant to the algorithm.
For the moment, automatic dialect recognition technologies are still in its pilot phase. Elaborating on Welfens’ contribution, in arguing that caseworkers’ security decisions are shaped by these sensing technologies is not to mean that they are going to be determined by it.7 As Scheel (2024) shows in his article, asylum administration in Germany is characterized by a reliance on a diverse array of techniques and technologies that shape the exercise of discretion. In this context, the implementation of the automated dialect recognition has, in effect, failed to capture the interest of caseworkers (Scheel, 2024: 8). The use of biometric sensing has been shown to frequently exceed the intentions of its designers, requiring constant “tinkering” to make it operational (Kloppenburg and Van Der Ploeg, 2020; Van Der Kist, 2024 see also Klonowska’s contribution). More importantly, migrants also have agency in resisting or circumventing these biometric control technologies (Scheel 2019). In short, not only do these examples illustrate how perceptions and actions are enabled and constrained by technologies (Gibson, 1979; Verbeek, 2005), they underline the need for situated analysis of whether and how these technologies are further taken up by their users or resisted by their subjects.
In extension, reflecting on the potential impacts of “sensing” in security studies should not just be concerned with the mediation of human perception by technological artifacts. I argue that it is equally, if not more important, to engage with what is rendered imperceptible by these machines (see also Brighenti, 2007). For instance, by carefully analyzing the affordances of sensory technologies like the automated dialect recognition system (amongst the many other technological devices that are currently being piloted and implemented) may bring into view the general shift in what can be said and heard in refugee status determination. This implicates the asylum administrator, whose perceptions are ordered by the mediation of the sensing technology, but even more so the asylum seeker, whose capacity to speak is reduced to a (meaningless) calculation (Lorenzini and Tazzioli, 2018). In short, by working in these different experiential registers, it will inevitably condition what might come to make sense in these administrative-legal domains: what bodies or voices can be rendered (in)visible, (un)sayable, or (in)audible (Rancière et al., 1999).
To this end, the conceptual and analytical potential of “sensing security” needs to be qualified further in terms of its political implications. For one thing, as the example of automated dialect recognition makes clear, its affordances echo a politics of control in which asylum seekers have been spoken for and silenced. In other words, the space for migrants to make their voices heard is, in part, limited by the technological constraints.8 Famously, Rancière et al. (1999) pointed to important moments in politics when those who “have no part” assert themselves as equal speaking beings—notably by interrupting the dominant perceptual regime. We may struggle to find these political moments in the examples of sensing security provided in this collective piece. But as I argue elsewhere (Van Der Kist, 2025), the problem with speech recognition technology is not only political but also legal: it revolves around the fact that claims to protection have traditionally been articulated through speech acts—mediated by the spoken or written word (Hildebrandt and Rouvroy, 2011). And it is precisely this ability to become audible, (visible and legible) as a political-legal subject, in close proximity to the sovereign power of the state, that is at stake in technologically-mediated practices of sensing security.
Concluding Thoughts
This collective discussion explored various security contexts from the perspective of “sensing.” Whether it was the detection of military threats, the targeting of enemies, the surveillance of students or the identification of migrants, our individual contributions shared a common concern with the ways in which these experiences of security are configured so that they can be acted upon. Despite our different case studies and approaches, the individual contributions converged on the mediated character of perception and action. While studies of security in IPS that concern themselves with human and non-human relations, knowledge production, and algorithmic governance share similar concerns, they were not the prime focus of our collective discussion. Rather, our contribution started from the perspective of how security is “sensed” and “made sense of.” It, therefore, builds on existing concerns in IPS scholarship whilst also shedding new light on the subject via its empirical and conceptual reorientation.
Accordingly, taking “sensing security” as a point of departure contributed to the discussion about what it means to make sense of security either as humans (with our own bodily senses) or how this is—fully or in part—delegated to technologies. Klonowska’s contribution on the “single pane of glass,” for example, highlights the mediating role of sensory technologies in extending military command-and-control of the battlefield. But while the (data-driven) technologies discussed may extend perceptions beyond what humans can do, Van der Maarel observes that human sensory capabilities still serve as a primary source of inspiration for the design of such technologies. In addition, the collective discussion underlines that the process of making security sensible and perceptible is never a neutral one. Most notably, Van de Ven’s analysis of “classroom sensors” has empirically detailed the particular dialectics of reduction and amplification at play in sensing security. Our contributions, therefore, show that these translations—or “transductions”—that make sense of certain aspects of the world to be secured often do so at the expense of other ways of experiencing, interpreting, and acting. Finally, by thinking through sensing security, the collective discussion has shed new light on the power dynamics at play in the will to govern security risks and threats, especially in the way perceptions are (re)distributed and (sovereign) decisions are enabled and constrained. Welfens’ contribution has offered a situated understanding of security perceptions and judgments, demonstrating that these are shaped by the interaction between frontline workers and “low-tech” bureaucratic tools. In extension, Van der Kist presented the argument that algorithmic technology has the potential to fundamentally reshape not only the perceptions and actions of asylum officials, but also the experience of asylum seekers as legal subjects.
This collective discussion piece does not claim to provide a comprehensive overview of sensing security. On the contrary. It is intended to stimulate further research and debate in IPS into the plethora of ways in which insecurity is perceived and acted upon. Such analyzes need not focus on the role of sensing in the context of governance. Further areas or themes of sensing security can also be explored. First, the relation between sensing security and political struggle, especially in relation to questions about which bodies or voices are rendered perceptible in security regimes and therefore political (Butler, 2005; Rancière et al., 1999). An inspiring line of research for IPS in this regard is collective or contrarian modes of sensing, for example, of citizens collecting sensory data to demonstrate (environmental) risks (Gabrys and Pritchard, 2018) or monitoring deaths at the (maritime) border (Heller and Pezzani, 2014). However, the binary between perceptibility/imperceptibility should also be problematized further in this context, in particular in relation to migrants that seek “to elude the gaze of dominant regimes of visibility altogether and strategically seek to remain imperceptible” (De Genova et al., 2022) (De Genova et al., 2022: 797). Finally, while most of our collective discussion piece speaks to current concerns in IPS about data and algorithmic governance, we feel that there is still more work to be done on this front from the perspective of sensing security. Most notably, investigations can proceed on the specificity of algorithmic sensing technologies, especially on the mediated sensing of fully autonomous technology.9
3 Algorithmic Security Vision
Diagrams of Computer Vision Politics
Authors: Ruben van de Ven, Ildikó Zonga Plájás, Chaeyuen Bae and Francesco Ragazzi.
This text is published in Environment and Planning D: Society and Space (2025): 10.1177/02637758251406481.
Note on inclusion in the dissertation: In accordance with PhD regulations regarding co-authored works, the primary responsibility for this article rests with me. As first author, I was responsible for the conceptualisation of the diagramming method and the development of its software. Furthermore, I led the interviews process, the subsequent analysis, and the writing of the manuscript.
More than ever before, security systems are using machine learning algorithms to process images and video feeds, in applications as diverse as facial recognition at the border, movement recognition in urban security settings, or emotion recognition in judicial proceedings. What is at stake in the technical and political transformations brought about by these sociotechnical developments? This article charts the development of a novel set of practices which we term ‘algorithmic security vision’ using diagramming-interviews as an exploratory method. Based on encounters with activists, computer scientists and security professionals, it identifies five interrelated shifts in security politics: the transition from a ‘photographic’ to a ‘cinematic vision’ in security; the emergence of synthetic data; the prominence of error—not as a defect, but as a central characteristic of algorithmic systems; the displacement of responsibility through reconfigurations of the human-in-the-loop; and finally, the fragmentation of accountability through the use of institutionalised benchmarks. Neither issue can be easily disentangled from the other; the study of algorithmic security vision thus unveils a rhizome of interrelated processes. As a diagram of research, algorithmic security vision invites security studies to go beyond a singular understanding of algorithmic politics and to think instead in terms of trajectories and pathways through situated algorithmic practices.
Introduction
In an increasing number of cities and at international borders, algorithms process streams of images produced by surveillance cameras. For decades, computer vision has been used to analyse security imagery using basic computation to, for example, send an alert when movement is detected in the frame, or when a perimeter is breached, based on the number of pixels changing colour. More recently, the increase in computing power and advances in (deep) machine learning is rapidly reshaping the capabilities of such security devices. They no longer simply quantify vast amounts of image sensor data but identify patterns within it to produce assessments and prompt interventions to a previously inconceivable degree. Pilot projects and off-the-shelf products are intended to distinguish individuals in a crowd, extract information from hours of video footage, gauge emotional states, identify potential weapons, discern normal from anomalous behaviour, and predict intentions that may pose a security threat. Security practices are substantially reconfigured through the use of machine learning-based computer vision, or what we call algorithmic security vision.
Algorithmic security vision represents a convergence of security practices and what Rebecca Uliasz calls algorithmic vision: the processing of images by machine learning techniques to produce a kind of ‘vision’ that makes realities actionable (Uliasz, 2020). It does not promise to eradicate human sense-making but rather allows a reconsideration of how human and nonhuman perception are interwoven with sociotechnical routines. Algorithmic security vision thus draws together actors, institutions, technologies, infrastructures, legislations, and sociotechnical imaginaries (see Bucher, 2018: 3). How does algorithmic security vision work—how does it draw together these entities—and what are the social and political implications of its use? While starting from specific technical developments, we are less concerned with the sole technical features of the systems than with their relation to the societal and political projects that are embedded in the technical choices made in their construction. This article therefore sets out to map out sociotechnical practices in which ‘algorithmic vision’ and ‘security’ practices feed into each other and explores how their entanglement reframes what it means to see, sense, surveil, and ultimately exert power.
To grasp the specificities of algorithmic security vision, we turn to professionals who work with those technologies. While we inevitably come to these conversations with preconceived notions and assumptions, we want to refrain from using pre-established, sedentary categories with which to map out these heterogeneous assemblages. Rather, we want to explore how the boundaries of algorithmic systems are drawn and negotiated, and how entities solidify and stabilise as they circulate between sites of development and deployment. To that effect, we introduce an inductive approach and methodological device: a time-based diagramming tool with which we combine interviews with a real-time drawing of diagrams. This setup allows us to engage with the contours and traces of algorithmic security vision and their coming together through an open-ended, associative and processual approach. Our interviews start with the simple question: can you draw the relation between computer vision and security for us?
In what follows, we begin by contextualising our exploratory theoretical and methodological approach within practices of mapping and diagramming. Then, drawing on our analysis of time-based diagrams, we identify and discuss five interrelated transformations in the security politics of algorithmic vision. First, we show the emergence of moving images, and the transition from what we define as ‘photographic’ to ‘cinematic vision’. Second, we describe and assess the implications of the emergence of synthetic data. Third, we acknowledge the prominence of error – not as a defect, but as a central characteristic of algorithmic systems. Fourth, we outline the reconfigurations of the human-in-the-loop dynamics. Fifth, we address the fragmentation of accountability resulting from the widespread use of benchmarks. Each of these empirical cases generates its own questions. In the conclusion, we reflect on how these come together to form a rhizomatic politics of algorithmic security vision.
Diagramming algorithmic security vision
The proliferation of algorithmic security systems has prompted extensive critical reflection across geography, science and technology studies, and critical security studies. Scholars have interrogated the anticipatory logics of algorithmic governance (Amoore, 2014), the rise of surveillance capitalism (Zuboff, 2019), the politics of operative images (Farocki, 2004), and the emergence of new forms of sensory and infrastructural power (Andrejevic and Burdon, 2014; Isin and Ruppert, 2020). Several authors have explored the use of algorithmic techniques in the treatment of images as security practices (Andersen, 2018; Bousquet, 2018; Fisher, 2018). They show how data, sensors, and algorithms are not neutral tools but are shaped by and constitutive of political logics. Importantly, much of this work recognises the entanglement between sensing infrastructures and algorithmic operations: What is captured, processed, and acted upon is conditioned by sociotechnical assemblages that cut across platforms, environments, and institutions.
However, while this entanglement between the micro and the macro dimensions is often acknowledged, it stays implicit and rarely made visible in tangible, empirical terms. The literature tends to analyse these relations at the macro level—focusing on systems of governance, platform capitalism, or border regimes—or track them at the micro level, analysing sites of implementation or precise algorithmic mechanisms. What is often missing is a means of articulating how these levels interrelate in practice and a way to find out which relations are most important. For instance, how is a local security system embedded in a neighbourhood shaped by and feeding back into broader data infrastructures, policy protocols, and institutional rationalities? What we point out is not a failure of theory, but a methodological challenge—how to capture and trace, with as few assumptions as possible, the specific nature, scale, and type of relationality that binds sensors, vision algorithms, administrative layers, and situated practices.
Diagrammatic mapping, we argue, offers a methodological approach capable of rendering such relations empirically while keeping the enquiry open. Its strength lies in its flexibility: it can capture both the situated micro-politics of specific systems and the macro-dynamics of the infrastructures, protocols, and institutions. It allows researchers and participants alike to foreground unexpected actors, neglected connections, and emergent logics.
We are not the first to critically examine algorithmic practices using diagrams. Computational practices are historically saturated with drawing and diagramming (Soon and Cox, 2020: 214). Kate Crawford and Vladan Joler teased out the many material facets of Amazon’s Echo device in their 2018 piece, the Anatomy of an AI System (2018). Joler and Matteo Pasquinelli mapped out the limits of artificial intelligence on a 2D surface (2020a). Perhaps most relevant for our aims is the work of Louise Drulhe, who in her Critical Atlas of Internet (2015) visually explores the politics of the various metaphors of the internet; this work does not fold all entities onto a single ordering logic but allows instead for the multiple competing renditions to co-exist.
As a method, diagramming does however something more. Across various intellectual traditions, including mathematics, logic, semiotics and philosophy, diagramming has been conceptualised not only a tool for representation and clarification, but as a way to enable a distinct ontological and epistemological approach to knowledge production (Stjernfelt, 2007). Diagrams enable us to ‘think differently’ about the world. According to Charles Sanders Peirce, diagrams allow indeed a distinct way of reasoning and knowledge discovery. What he conceptualised as ‘skeletal icons’ (Stjernfelt, 2011) makes relational relations explicit, contrasting with the limited value of abstract statements without their aid (Peirce, 1931–1966). For Peirce, a map, often considered a quintessential diagrammatic form, is a semiotic subtype of the diagram, sharing an identity while potentially incorporating pictorial elements (Gerner, 2010). Diagrams can represent relational structures and facilitate deductive reasoning through observation and manipulation. These initial intuitions were formalised by the cognitive science perspective in the work of Jill Larkin and Herbert Simon. Larkin and Simon distinguish ‘sentential’ forms of representation (sequential, argument-based forms of knowledge) to diagrammatic representations (relational, multidirectional, and networked information). They show that contrary to the former type of information, the latter needs to be visualised spatially in the mind before it can be assimilated and understood (for example, an organisation’s reporting organigram, or an electrical circuit). Diagramming on paper or on a screen is thus a way to lighten the mental load on the brain, which would otherwise need to draw the diagram for itself, mentally, to understand complex relations or processes (Larkin and Simon, 1987).
While these approaches highlight key dimensions of diagramming as cognitive and epistemic practice, they may at times assume that the diagram is in a passive relation to the mental image—a representation of what needs to be understood or communicated. The work of Deleuze, however, points out that the diagram is also a device that ‘does not reproduce the visible but constructs the conditions of visibility itself’ (Deleuze, 1988: 34). For Deleuze, diagrams, such as the figure of the ‘rhizome’, are indeed ‘abstract machines’ that map forces, intensities, and becomings, functioning within a plane of immanence that precedes and destabilises conventional signification (Deleuze and Guattari, 1980). Unlike Peirce, for whom the map is a subtype of the diagram, Deleuze suggests that a ‘cartographic ontology’ (Gerner, 2010: 100) precedes the diagram category itself. The diagram, in this view, is a ‘map of virtualities, superimposed onto a real map, whose distances [parcours] it transforms’ (Deleuze, 1994; cited in Gerner, 2010: 55, 100). Diagrams thus actively coproduce knowledge of the sociotechnical practices they are to represent.
Drawing on these contributions, we identify three key affordances of diagramming that make it particularly suited for analysing algorithmic security: (1) relational thinking; (2) manipulation of abstract relations; and (3) generative open-endedness (Table 1).
| Affordance | Enabling characteristic | Description |
|---|---|---|
| 1. Relational thinking | Spatialisation | Diagramming translates thought into spatial form, allowing elements to be perceived simultaneously in complex relational networks rather than in a series of linear sequences. The diagrams’ ‘speculative geometries’ (Soon and Cox, 2020: 221) foreground multiplicity, proximity, contrast, or absence—qualities that help bring implicit hierarchies, associations, or gaps to the surface. |
| 2. Manipulation of abstractions | Graphical inscription | Hand-drawn diagrams, with their ‘sketchy’ graphical form are neither definitive nor exhaustive renderings but facilitate manipulation and revision. Elements can be added, reconnected, or reformulated in response to new insights, enabling what Crilly et al. (2006) call ‘graphic ideation’: an iterative process in which the interviewee continuously tests their visual expression of ideas (Crilly et al., 2006; McKim, 1980), bringing an extra layer of critical reflection (Bravington and King, 2019; Hurley and Novick, 2006). |
| 3. Generative open-endedness | Multiple layering of meaning | Diagrams operate on multiple semantic levels, combining denotation and connotation, allowing for interpretive ambiguity. This makes them, as O’Sullivan (2016: 13) notes, ‘protocols for a possible practice’—open-ended, exploratory, and generative. It is precisely the uncertain status of what appears on paper — being both an ephemeral idea a definitive communication — that sparks the conversation (Bagnoli, 2009; Crilly et al., 2006). |
In our approach, we thus mobilise these three characteristics of diagramming—spatial thinking, graphical manipulability and multiple layering of meaning—as tools for both data collection and elicitation, emphasising not only the final output—a static image—but also the thought processes embedded in its production.
Methodology
For this article, we conducted eleven unstructured interviews with twelve professionals in computer vision in the field of security. The participants were purposively recruited based on their direct experience of and involvement in at least one of the following activities: (1) the design and development of computer vision algorithms and models; (2) the practical deployment, operational integration, and routine management of these systems in real-world security contexts; or (3) active critique, resistance, or contestation at a policy level of the implementations of computer vision technologies in contexts of (in)security.
Given practical constraints such as the specialised nature of this field and the limited pool of experts, we utilised convenience and snowball sampling techniques to recruit interviewees, primarily guided by participant availability, ease of access, and strategic relevance to the research question. Our selected interviewees should not be understood as statistically representative of a broader population or as ‘illustrative representatives’ of the field (Mol and Law, 2002: 16–17). Rather, through purposive sampling, interviewees were deliberately chosen to ensure variation across cultural and institutional affiliations, professional roles, and types of engagement with algorithmic practices.
The interviews were conducted in English in the Netherlands (6), Hungary (3), Germany (1) and Poland (1) by Ruben van de Ven and Ildikó Plájás, and subsequently by Clemens Baier. We employed a predominantly unstructured interviewing approach, initiating each conversation with two guiding prompts aimed at eliciting initial reflections. We began by asking the interviewees if they use diagrams in their daily practice. We then posed the prompt: ‘When we speak of ’security vision’, we speak of the use of computer vision in a security context. Can you explain, from your perspective, what these concepts mean and how they come together?’ Initial interviews revealed emergent thematic patterns, prompting the adaptation and refinement of subsequent interview questions to explicitly address these recurring themes.
At the practical level, we presented our interviewees with a large, A3-sized, digital tablet, and asked them to draw a diagram while answering our questions. Ruben van de Ven programmed an interface that could record conversation in drawing and audio. The participants could not delete or change their drawings, so their hesitations and corrections were preserved.10 Breaks appear in the conversation as the interviewees need to think about how to draw (Bravington and King, 2019: 509). Sometimes, the participants seemed uncomfortable about their ability to draw figuratively (see also Copeland and Agosto, 2012). These moments allowed us to clarify that our interest is not so much in the figurative as much as in the relations they drew. Some of the participants therefore opted to diagram with words, rather than figures.
Some of the diagrams that were drawn during the conversation show similarity with the kind of figures the practitioners work with in their own practices. Diagramming is indeed a key method in the field of technology, most notably in the conceptualisation and design of computational practices (Mackenzie, 2017). This can primarily be seen in the use of flow diagrams (Soon and Cox, 2020: 221), that represent different states or stages of computational processing as data ‘travels’ through a system. In our conversations, this style of drawing is particularly visible with those who develop or manage systems for algorithmic security vision. One interviewee explicitly mentioned he was reproducing his slideshow presentation in drawing as he went to explain some of the key concerns of his project. Many of the figures that appear in drawing thus do not exist purely for the conversation, but resemble or reproduce, both in form and content, the figures used by the practitioners in their daily practices. However, as the interviewees draw them on the tablet, they invite reflections that provided crucial insights. For example, one interviewee, a legal scholar and activist, drew a face-database, after which he reconsidered the visual representation: face-databases do not necessarily contain images of faces, but facial features in numeric form. While he initially drew on a common cultural representation of face recognition and surveillance, his technical expertise made him go against such a way of visualising these technologies (van de Ven and Plájás, 2022 in this bundle as Annex C). Thus, the interviewees tweak the images by drawing new relations, adding layers of explanation, and, as the conversation progresses they refer back to what they drew before, ‘hyperlinking’ the conversation back to earlier concerns.
In the phase of data analysis, the software we developed allows the diagrams to be annotated, creating short video clips, facilitating comparisons among the conversations. All participants except two provided explicit consent for identification by name11. The multimodal tool was created and used to explore the possibilities offered by diagrams to elicit the main three affordances described earlier : (1) relational thinking; (2) graphical manipulation of abstraction; and (3) generative open-endedness in discussing the heterogeneous relations that constitute algorithmic security vision (see van de Ven and Plájás, 2022 for a more detailed account of the methodology and its affordances). Integrating temporal and dynamic aspects of diagrammatic practices was particularly significant for capturing not merely static visual representations, but also the performative processes through which these relational features emerge and are enacted. It is thus important to note that while the PDF version of this article contains only static image representations of the diagrams, the multimodal version allows accessing the full audio-visual excerpts of the diagrams12. Our analysis refers to the audio-visual material.
Diagrams of algorithmic security vision
In what follows, we traverse these diagrams and highlight five features of camera-based algorithmic security practices that help us to rethink some central notions of the literature on algorithmic security (Figure 1).
1. From photographic to cinematic vision
The first finding of our diagrammatic interviews is a noticeable shift in the temporal dimension of the uses of computer vision in security technologies. With the advancements in algorithmic models and computational power, more systems than ever are able to assist security decisions based on video data. This shift from still images to video has implied an exponential growth in data which necessitates advanced infrastructure and analytical techniques capable of handling and interpreting continuous streams of visual information. As Gerwin van der Lugt notes, ‘Most cameras do about 30 frames a second. So if you have for example [in] Amsterdam 300 cameras [with] 30 frames per second for what is it? This is basically the number of frames generated each week. Yes, about 100 million’.
This, in our view, marks a shift from what we could call a ‘photographic’ mode of algorithmic analysis, such as biometric imaging (Pugliese, 2010), facial, iris and fingerprint recognition (Møhl, 2021), and body scanners (Leese, 2015). While the former type of systems works on the premise that it can identify individuals on the basis of unique and immutable features of their body (a fingerprint, an eye iris), the new, ‘cinematic’ systems work instead from the assumption that temporary anomalous states, grounded in suspicious movements or emotions, for instance, are what allows the identification of suspects, regardless of their intrinsic features. In the ‘cinematic’ mode of analysis, suspicion is thus not established only inside a single frame, but across a succession of them. This does not mean that both logics cannot be used simultaneously within an integrated system (for example looking in a crowd at movements of people previously identified through facial recognition), but rather that the ‘cinematic’ mode is a relatively new feature that comes with its own affordances and politics13.
Two diagrams represent this shift particularly well. Figure 2, by Ádám Remport14, illustrates the photographic mode. It depicts how photo snapshots produced by several surveillance cameras equipped with face recognition technology analysed over a longer period of time, can identify the same person visiting a bar, a church, an NGO. Remport spatialises the functioning of a facial recognition system, which constructs movement through the placing of cameras along the route of an observed individual. Here, the diagram quite literally maps the ways in which a series of static data points (images) can be used to infer and analyse movements and life habits over time and space.
Figure 3, by Gerwin van der Lugt15, shows how the alternative is ontologically different, through a different graphical spatialisation of the relations between the same entities. The system van der Lugt develops is meant to analyse the behaviour of people by looking specifically at movements and distinguishing different behavioural patterns. This system, as he explains, ‘can just detect things like punching, hitting, kicking, dragging someone or pushing someone and then combine those detections and at some point conclude that there is an ongoing violent incident or aggressive behaviour’.
The drawing on the top of Figure 3 represents the photographic paradigm: when processing video data, the frames (f1, f2, f3) are passed to an object detector (YOLO), which runs independently on each frame, one after the other. In the second diagram, the video frames are no longer processed as individual frames but coalesced into a single data unit of 10 frames (f1–f10) that is passed to the algorithmic model (M). Multiple 10-frame units are then analysed by an additional process that looks at how these groups of frames relate to each other. This allows models to identify and categorise patterns across frames that happen in time, for example a hand gesture. By using the diagrammatic schemata to highlight the distinction between the two systems, the diagrams reveal that it is not a variation of the previous, photographic logic, but a profoundly different one.
This shift has political consequences in how models perform suspicion. The cinematic logic no longer relies on the inscription of supposed ‘indelible’ truths in biological features of individuals (faces, fingerprints, retinas); instead it operates by categorising unfolding movement as suspicion (a suspicious gesture, gait, facial expression), encoded across the frames. The advent of the cinematic dimension thus constructs suspicion as time-based, ephemeral and changing. Suspects no longer need to be categorically present in or absent from a space—it does not matter if they enter the NGO office—but bodies are marked by their movements through space, they can be passing, traversing, entering, or escaping. Thus, while earlier studies on biometric technologies have located the operational logic on identification, verification, authentication, in other words, knowing the individual (Ajana, 2013; Muller, 2010), the cinematic vision finds its operational logic in the mobility of embodied life (see Huysmans, 2022), and its governance through the dynamism of ever-changing ‘real-time’ clusters (Isin and Ruppert, 2020).
2. Synthetic data
A second major theme that emerged from our diagram-interviews is the growing significance of synthetic data in reshaping AI training pipelines which raises new ethical concerns. Over the past two decades, much of the critique of algorithmic security systems has centred on the tension between the ‘reality’ they claim to capture and the processes through which those realities are rendered computable. Central to this critique is Haggerty and Ericson’s notion of the ‘data double’ (2000), which highlights the constructed, arbitrary and always imperfect nature of data representations. This process of translation is indeed never neutral; it involves choices about what to include, what to exclude, and how to categorise, fundamentally shaping the reality that the system perceives and acts upon. Once encoded, these data profiles become the basis for social sorting (Gandy, 2021; Lyon, 2003), where realities are increasingly shaped by data-driven inferences and interventions, potentially leading to new forms of normativity and control (Bigo et al., 2019). The data used to train self-learning systems—which functions as ‘ground-truth’—is thus never a neutral or ‘raw’ representation of the world (Gitelman, 2013), but rather a product of selective encoding that risks perpetuating societal biases and representational shortcuts (Malevé, 2020).
Our research shows that synthetic data complicates these debates even more, appearing to actors both as a possible remedy and a further destabilising element to the question of representation of reality. In Figure 4, Sergei Miliaev16 describes the process through which an AI model is trained. The cycle begins with data acquisition, signalled by the words ‘data’, followed by ‘annotation’, ‘prototyping’, ‘deployment’, and an arrow back to ‘prototype; to emphasise the feedback-based refinements. The AI models are typically using real-world datasets, explains Miliaev in the interview, to generate inferences which can be acquired through ’web scraping’, by relying on ‘partner data’, and open-source footage (see the upper left corner of Figure 4).
Yet sourcing such real-world datasets—as shown in the controversies around large language models (LLMs) and copyright issues (Cyphert, 2024)—is a key challenge for the AI industry. Miliaev’s diagram allows us to pinpoint that this issue emerges from the very beginning of the AI training workflow; while hidden behind a sequence of processing steps—indicated by arrows in the diagram—the deployment of algorithmic security vision relies on a range of data sources, all of which raise particular concerns. Acquiring ethically sourced, high-quality data—especially for scenarios that require very specific data, like video—is increasingly difficult due to privacy laws and annotation costs. Moreover, these datasets—when they exist—either contain very little data, are of bad quality, or do not adhere to the requirements of the model.
To circumvent the challenges posed by the availability of reliable, varied and open access data Miliaev adds, thus introduces, another data source in the lower left corner of his diagram: synthetic data. The diagrammatic format allows to easily lay out the structure of relations graphically, evaluate it, and insert a new node (synthetic data) within the existing configuration. Synthetic data, which consists of computer-generated images created via 3D modelling, or by means of generative AI (genAI), can be a useful ‘workaround’, so explains Miliaev. Synthetic data enables the creation of large, labelled datasets tailored to specific scenarios, from suspicious activity to pre-crime behaviours. Crucially, because it is generated rather than captured, synthetic data bypasses some of the most stringent privacy restrictions. Furthermore, manipulated data can, for example, simulate diverse conditions—lighting, actions, faces—that are otherwise hard to capture at scale. Miliaev explicitly cites Microsoft’s work on training facial recognition models entirely on synthetic faces (Bae et al., 2023) as proof of this potential.
The adoption of synthetic data thus forces us to problematise the politics of ‘ground truthing’, i.e. the practice of establishing what the models should consider as ‘true’. One is fidelity: the challenge of ensuring synthetic scenes accurately reflect the variability and messiness of real-world environments. Poorly generated data—lacking realism in lighting, weather, or human behaviour—can lead to models that fail when deployed. Another issue is identity representation. While synthetic datasets allow designers to manipulate ethnicity, gender, age, and other attributes in a hope to de-bias the collection, Miliaev notes the difficulty of modelling long-term identity changes like ageing or behavioural variability. If the parameters used to generate this data are biased, the resulting models will inherit and even amplify social biases, as synthetic data does not escape the representational politics that shape captured data.
The diagram thus reveals how synthetic data destabilises conventional epistemologies. While conventional data sources for machine learning, such as operational data or web-scaping, have long been considered to provide access to a ‘raw’ or unfiltered truth, years of critique have established that data are never raw; they are constructed, interpreted, and made intelligible through a relational network of institutional, technical, and epistemic filters (e.g. Gitelman, 2013). This critique has now been taken up by many who develop algorithmic technologies, presenting itself as a problem in need of a technical solution. To this end, Miliaev annotates synthetic data with a moral assessment—‘fairness’—which is to be obtained by balancing various ‘attributes’ of the faces in the dataset different from what would be found in previously considered raw data. The new flow diagram that emerges no longer consists of merely technical description, but becomes indicative of moral and ethical assessments that lie at the heart of data gathering and processing. In what way, however, is this fairness established? Does the introduction of these debates imbue the diagram with ethical considerations, or is the ethics rendered as a technical concern? The ‘mean images’ of synthetic data—visual renderings of a non-existent, probable median—reproduce normative conceptions of reality, effectively complexifying the claims to indexicality of lens-based images (Steyerl, 2023). Instead of referencing empirical events, synthetic datasets rely on speculative associations and probabilistic reasoning, reframing our understanding of evidence, ‘ground truth,’ and credibility in machine learning practices.
3. Managing error: from the sublime to the risky algorithm
The question of the ‘ground truth’ is indissociable from the third central theme emerging from the diagrammatic engagements: the constitutive role of error in algorithmic security systems. While much literature addresses algorithmic prediction through self-learning systems (Azar et al., 2021), most of it has focussed on how these systems produce risky subjects (Amicelle et al., 2015; Amoore and De Goede, 2005; Aradau et al., 2008; Aradau and Blanke, 2018), based on environmental and individual features (Calhoun, 2023). In this critical body of literature, practitioners working with algorithmic technologies are often critiqued for understanding software as ‘sublime’ (e.g. Wilcox, 2017: 3), meaning that they are thought to believe that their systems ‘work’ in what they claim to do.
The diagrams we collected force us to rethink this premise. Models cannot be assessed through a binary criterion—working/not working – they are instead assessed through a series of error metrics. This is not a flaw that actors aim to fix once and for all; it is the asymptotic quest towards a complex mix of lowest error rates that drives the work of algorithmic security vision. Our diagrammatic interviews show both the importance of the error rates, but also the complexity to account for them and explain precisely what they are.
Figure 5, by CTO Gerwin van der Lugt is particularly enlightening in this regard. As he explains, the most prominent way in which error figures in model development is in its quantified form of the true positive rate (TPR) and false positive rate (FPR). Van der Lugt stresses the significance and definition of these metrics by marking them as the key variables to optimise, on the top right corner of the page (Diagram 5). Van der Lugt initially describes the FPR is the number of false positive classifications (FP) relative to the ‘input’, the number of video frames being analysed. When he continues to explain the TPR, he similarly marks it on the tablet as being equal to the number of true positive classifications (TP) over all inputs. As he goes on however, he realises that this is slightly incorrect. In fact, it should be ‘true incidents’. Looking back at the tablet again, he adjusts the original FPR equation: it should be specified as FP over the number of ‘false inputs’ (and not just all inputs). ‘These are quite tough to define but they are quite important, as these need to be really low’. The diagrams work here as an important support for van der Lugt to manipulate the abstract notions of errors through graphical inscription: by first noting them, then correcting them, the diagramming helps refine the explanation of a key yet complex notion to explain, even for the developer of the system. Moreover, as he corrects his definition, van der Lugt comments not only on the difficulty of these definitions, but also on the importance of having them right, as the these definitions determine the work of his development team, the ways in which his security operators engage with the technology, and how the system is evaluated by his clients.
As he goes on, van der Lugt ponders that algorithmic security vision is inherently prone to error, which affects the implementation of security practices. He illustrates this with Oddity.ai’s violence detection system, questioning whether playful fighting (stoeien) should be classified as violence. He argues for distinguishing between false positives—errors in evaluation—and errors in how the security problem is operationalised. He offers two reasons. The first reason is that excluding stoeien could reduce the algorithm’s TPR, disrupting the balance between TPR, FPR, and other performance metrics. The developers aim for fewer than 100 false positives per 100 million frames weekly, since too many alerts can desensitise human operators. The second reason is that excluding stoeien might introduce subtle biases—for instance, the algorithm might infer violence based on age rather than behaviour. Van der Lugt warns that such discrimination is both undesirable and hard to detect. In this view, error is not just mathematical but something to pre-emptively manage, placing responsibility on developers to anticipate and mitigate failure.
Another aspect concerns the mutable nature of threats and susceptibility to adversarial attacks. András Lukács17 notes that detection models ‘can be cheated by quite simple tricks’. Even data-rich systems remain vulnerable to intentional evasion, highlighting the challenge of maintaining relevance in evolving threat landscapes. John Riemen18 adds that poor input quality undermines facial recognition—for instance, the wrong still from a video may fail to identify correctly.
Practitioners suggest several strategies to respond to these issues. One is narrowing the algorithm’s scope to high-impact, clearly defined crimes. Guido Delver19 recounts moving from detecting vague ‘suspicious behaviour’ to focusing on burglary. Van der Lugt’s firm similarly focuses on weapons, vandalism, and physical violence. This narrowing reduces ambiguous boundary cases—both an ontological and an operational shift. Another strategy is continuous retraining. Dirk Herzbach20 explains how operators feed annotated alerts back into the system. Parameter tuning to balance false positives and negatives is also key. This balance is political: as Jeroen van Rest21 notes, communities vary in their tolerance for error—some may accept high false positives for safety, others may not. Thus, the acceptable error threshold becomes a negotiation site between developers, publics, and institutions.
In this light, error is not external to algorithmic systems—it is their condition of possibility. As Pasquinelli (2019) puts it, machine learning is based on ‘formulas for error correction’; error is not a failure to remove, but the mechanism through which learning occurs. Amoore (2019) similarly notes that ‘it is precisely through these variations that the algorithm learns what to do.’ Error, initially a mathematical concept, permeates every discussion of algorithmic security vision. What emerges is a double articulation of risk. While risk is usually treated as external to, and produced by, security technologies, technologies of algorithmic security vision themselves become objects of risk management. The inevitable errors of machine learning algorithms place risk not outside, but at the core of the security apparatus.
4. Reconfiguring the human-in-the-loop
The understanding of algorithmic security vision as a practice of managing error has significant consequences for the role of the human within these assemblages—our fourth main theme. The critique of algorithmic security has often construed the human-in-the-loop as one of the last lines of defence to the inevitable erroneous outcomes of automated systems (Markoff, 2020). However, critical security studies have questioned this representation, emphasising the visualisation of algorithmic predictions in graphical interfaces (Aradau and Blanke, 2018) and how the operator's embodied decision-making is intertwined with the algorithmic system of which they are part (Wilcox, 2017). Furthermore, operators may find themselves uncertain about the system's functioning when errors occur (Møhl, 2021). In other words, when considering questions of responsibility, a system's operator cannot be seen as distinct from the algorithmic assemblage (cf. Hoijtink and Leese, 2019a). The question in a system's design thus is not whether a human operator makes autonomous decisions but rather, how they are to interface algorithmic predictions, and how issues of agency and responsibility are negotiated. The diagrams illustrate how this leads to differentiated processes of (in)visibilisation.
Looking across the various diagrams reveals a spectrum of system designs concerning the autonomy of security operators and the technical expertise they are expected to bear. In a first kind of designs, the human operator is central to the decision-making processes, acting as the interface between, and external to, algorithmic systems and surveillance practices. For example, Dirk Herzbach explains that when the Mannheim police is alerted to an incident by the system, it is the operator who decides whether to deploy a patrol car. Here, the human-in-the-loop has full agency and responsibility for operating the (in)security assemblage, with the capacity to evaluate and utilise algorithmic systems selectively and with care. Herzbach therefore prefers an operator to have patrol experience, so they can best assess which situations require intervention. He is concerned that knowledge about algorithmic biases might interfere with such decisions. In other words, the security operator is considered the expert of the subject of security and is expected to make decisions independently from the information that the algorithmic system provides.
However, in a second kind of design, the human operator is considered an integral part of the algorithmic security vision system, influencing and influenced by its operation. Such designs acknowledge that it might be more complex for the human-in-the-loop to perform as the primary counter to algorithmic error. This is illustrated by Guido Delver, project manager for the Burglary-Free Neighbourhood project in Rotterdam, which incorporates autonomous systems into street lamps. In Figure 6, Delver maps out the different ‘stakeholders’ of the project on a flat canvas, using the diagram as a way to spatialise relations of a heterogeneous set of key entities. He envisions the neighbourhood as a site where the public and private space—the residents and government—meet, and which also involves suppliers of technology and research institutions. By rethinking the traditional top-down hierarchy of the surveillance apparatus, Delver aims to counter government hegemony. Delver uses this rendition of the various stakeholders to ask the question ‘who [of these parties] owns the data?’. However, upsetting the traditional hierarchies comes with its own risks. Delver illustrates his point with a scenario in which the algorithmic signalling of a potential burglar may have dangerous consequences: ‘Does it evoke the wrong behaviour from the citizen? [They could] go out with a bat and look for the guy who has done nothing [because] it was a false positive.’ In this case, the worry is that the erroneous predictions will not be questioned. Thus, while often deemed to be ultimately responsible for the judgement that follows from the algorithmic process, human participation or ‘interference’ in the operation can likewise be rendered as harmful.
As a means to mitigate these concerns, in Delver’s project, the goal was to actualise an automated decision system (ADS), ‘with as little interference as possible’. Placed in the centre of his drawing, different stakeholders interact with this system in different ways. Thus, figuring the human-in-the-loop—whether police officer or neighbourhood resident—as risky can lead to the relegation of direct human intervention. For Delver, access to the data—both the data that makes inferences possible, as well as the outcomes produced by the device—needs to be curated.
The curation of data flows—acknowledging the fallibility of human decision making—is even more prominent in another diagram. In Figure 7, John Riemen (dis)assembles the process of facial recognition as it is employed by the Dutch police. After drawing three human figures, he stresses the relations among these multiple humans-in-the-loop, and how some data should be allowed to pass (indicated by the arrows in the diagram; for example, from a database expert who performs the search to an investigator of the case), while other flows of information should be blocked. For facial detection to be accurate, the human interaction with the algorithm needs to be sanitised of personal information.
By drawing the process as series of actors, all of whom are biased, Riemen sketches out how they are embedded in, and thereby influenced by, a set of mutual relations that needs to be regulated, here the diagrams capture the relational and sequential nature of these links.
The diagrams thus challenge the simplistic view of the human-in-the-loop as a fail-safe against algorithmic errors by highlighting the complex network of sociotechnical relations they are enmeshed in. For both Delver and Riemen, the inherent error—whether performed by a person or a computer—has to be mitigated by curating the transfer of information from one node to the next. More than multiplying the number of humans-in-the-loop, these diagrams encourage us to think in terms of ‘curatorial pipelines’ (Malevé, 2023) that mediate the different human and algorithmic operations.
5. Delegating accountability to benchmarks
The final theme of our diagrams concerns the question of accountability. Literature on the ethical and political effects of algorithmic vision has notoriously focussed on the distribution of errors, raising questions of ethnic and racial bias (e.g. Buolamwini and Gebru, 2018). This critique has now been internalised in the field of AI development; virtually all of our correspondents acknowledge the concern with bias.
To mitigate it, many AI developers—including some of our respondents—have come to rely on an external reference against which the error is measured: benchmarks. John Riemen, for example, who is responsible for the use of forensic facial recognition technology at the Centre for Biometrics of the Dutch police, describes how their choice of software is driven by a public tender that demands a ‘top-10’ score on a benchmark for facial recognition vendors, maintained by the American National Institute of Standards and Technology. Often shortened to the acronym NIST (e.g., Figure 7), this benchmark ranks facial recognition technologies of different companies by their error metric across groups. The mitigation of bias is thus outsourced to an external, and in this case US-based, institution.
By examining where John Riemen’s diagram ends (see Figure 7)—that is, where entities are no longer broken up into its component parts, but rather bracketed and considered as a whole—we get a sense of how issues are ‘black boxed’. As Riemen describes the process of facial recognition, he is able to explain and draw it in much detail. Therefore, when asked about bias mitigations, he is able to precisely locate the various measures of filtering that take place, drawing the ‘walls’ in between the different people (see the vertical line in the diagram). On the lower part of the drawing, the NIST database does not have the same detail in process, nor scrutiny. In Riemen’s diagram it becomes a single entity, represented by an acronym. As a benchmark, the facial recognition vendor test is implicitly assumed to have already implemented the measures that are required for its use by the Dutch police. The diagrams here provide a graphical illustration of the abstraction and erasure of particular processes, showing only the inputs and outputs within a general network. Put in more stark terms, where the diagram stops, the respondent’s responsibility stops. With NIST, the prevention of algorithmic bias comes to rely on a single benchmark: a de-facto standard that is managed by an external institution.
The problem is, while a particular kind of algorithmic bias (i.e., ‘demographic differentials’) is rendered central to the NIST benchmark, the mobilisation of this reference obfuscates questions on how that metric was achieved. For example, the NIST benchmark datasets are known to include faces of wounded people (Keyes, 2019). Moreover, questions about training data are invisibilised, even though that data is a known site of contestation; the Clearview company is known to use images scraped illegally from social media, and IBM uses a dataset that is likely in violation of European GDPR legislation (Bommasani et al., 2022: 154). Pasquinelli (2019) argues that machine learning models ultimately act as data compressors: enfolding and operationalising imagery of which the terms of acquisition are invisibilised.
This attention to the invisibilisation of and in/between processes reveals a discrepancy between the developers and the implementers of these technologies. On the one hand, the developers we interviewed expressed concerns about how their training data is constituted to gain a maximum FPR/TPR ratio; while showing concern for the legality of the data they use to train their algorithms—propelling for example the use of generative AI (see section 2). On the other hand, questions about the constitution of the dataset have been largely absent from our conversations with those who implement software that relies on models trained with such data. Occasionally this knowledge was considered part of the developers' intellectual property that had to be kept a trade secret. In such cases, for the implementers, a high score on the benchmark is enough to pass questions of fairness and to forgo any further inquiry into the inequalities propagated by the algorithmic system and the data through which it learns, thus legitimising its use. While the deployment of algorithmic models indirectly relies on the training data, it is not deemed relevant in the consideration for one particular model over another.
The configuration of algorithmic vision’s bias across a complex network of fragmented locations and actors, from the dataset, to the algorithm, to the benchmark institution reveals the selective processes of (in)visibilisation. It illustrates well how accountability is bracketed through the invisibilisation of the dataset as it is ‘compressed’, in Pasquinelli’s terms, into an algorithmic detection model, with the formalisation of guiding metrics into a benchmark. One does not need to know how outcomes are produced, as long as the benchmarks are in order.
Conclusion: a diagrammatic research agenda
In this conclusion, we reflect upon a final dimension of time-based diagramming as a technique for elicitation in, and analysis of interviews. What started as an endeavour to map out the politics of algorithmic security vision yielded something quite different. Traditionally, maps subordinate entities to a single overarching order. With diagramming, however, what emerges, instead of a singular rationale of algorithmic security vision, is a series of interrelated configurations. While writing this text, indeed, the search for a coherent structure through which we could map the problems that emerged from analysing the diagrams in a straightforward narrative proved elusive. It became evident that through the spatialisation of the conversations, the diagrams yielded a rhizome of interrelated problems (Figure 8).
Making sense of the diagrams, then, can perhaps be best compared to walking, which, Certeau (1984) reminds us, is a way of traversing space that does not provide an overview. While maps produce such overview by laying out entities as distinct, walking happens along trajectories that join things together (see also Mol and Law, 2002: 16). Rather than decomposing algorithmic security vision into distinct elements—algorithms, institutions, datasets—we cut across the diagrams, creating a multitude of possible inquiries and overlapping trajectories that bring issues into relation. When navigating the interstices opened up by our empirical research, we find a diagram of interrelated questions.
One such trajectory revolves around the question of suspicion. As systems shift from photographic recognition to the analysis of movement and behaviour, suspicion becomes dynamic—triggered not just by who is present, but by when and how they move. This shift propels other developments. Analysing movement instead of still images requires much more training data; data that is often not readily available. While facial recognition algorithms could be trained and operated on quickly repurposed photographic datasets of national identity cards or drivers’ licence registries, no dataset for moving bodies has been available to be repurposed by states or corporations. Therefore, cinematic algorithmic vision can be seen as one of the causes behind the increasing use of synthetic data to train models against rarely documented threats. By simulating potential violence or suspicious actions, these datasets encode imaginaries of threat. These imaginaries, translated into algorithmic systems, shape how security is enacted in practice—well before any real event occurs. This shift in the construction of suspicion opens a cascade of questions. Cinematic vision complicates the study on what gestures or trajectories prompt classification as a potential threat. As cinematic vision propels the use of synthetic data, it drives a whole different set of questions: who defines what scenarios are worth simulating? What data counts as realistic? These concerns reveal the entanglement of temporal analysis, synthetic data generation, and the configuration of accountability and responsibility, offering critical entry points for investigating how deviance is constructed algorithmically.
Closely linked to the use of synthetic data we can trace the issue of error and fairness of algorithmic systems. Virtually all of our correspondents acknowledge error as inherent to algorithmic security vision: there is no ‘sublime’ algorithm. Thus while mistakes are inevitable, the way the algorithm errs — and who is affected by these errors — is considered a design challenge. Every algorithm includes thresholds that define acceptable levels of error, whether false positives, false negatives, or biased outcomes. But who sets these thresholds, and on what grounds? What level of misclassification is deemed tolerable, and for whom? And how are harms traced when models misidentify, exclude, or discriminate? The negotiation of these boundaries involves not just developers but also policy-makers, system integrators, and institutional end-users. Some turn to ‘synthetic data’ as a way to mitigate the reproduction of social biases that stem from these errors. Others’ focus on the ‘curatorial pipeline’ of data—they consider which data is allowed to pass from one operator to the other—raises the question if there is such a thing as non-synthetic data (see also Gitelman, 2013). In many cases, measures are put in place to assess the fairness of an algorithm’s operation with benchmarks such as the NIST face recognition vendor test—inadvertently shifting attention away from questions around the fairness of what goes into an algorithm for its training. These decisions reflect broader power dynamics and ethical trade-offs. Understanding these structures is essential to assess how control is exercised, delegated, or contested within security infrastructures.
This brings us to a third trajectory we can draw through our diagram-based conversations: concerns on acceptability and accountability. For if systems err, the ideal of ‘human-in-the-loop’ oversight is often invoked to guarantee responsibility, while other cases—such as the Burglary Free Neighbourhood or the use of facial recognition at the Dutch police—suggest that the role of the human operator heavily depends on how their interaction with the algorithmic system is mediated; in some cases the presence of a human is considered a risk. Exploring the transparency and explainability of these hybrid systems, not just for the public but also for the human operators themselves, is crucial. Further research should thus critically analyse the ways in which accountability is architectured as humans, algorithms, and their interfaces jointly contribute to security decisions.
Together, these research questions form a comprehensive agenda for investigating the political, ethical, and social dimensions of algorithmic security vision. While one issue contains echoes of the other, the network of interrelated problematisations cannot be flattened into a single narrative. The constraints imposed by the linear structure of a text certainly necessitate a specific ordering of sections. Yet the different research directions we highlight form something else. The diagrams do not represent a map of how things are, but a drawing together generative of fresh tensions. These pathways link each question to overarching themes, thereby offering a diagram of research for future studies that seek to unpack the complex ways in which algorithmic practices are reshaping the contemporary security landscape.
4 Prediction of the now
Algorithmic surprise in the governance of movement
This text is currently under review at Big Data & Society
Camera surveillance in public space increasingly relies on algorithms that process streams of images to track the movements of bodies. This article contends that these security technologies demand a new account of the temporal logics by which movement is governed. Examining common accounts of algorithmic security politics, the paper suggests that both the immediacy of ‘real-time’ governance, and anticipation of future harm of pre-emptive governance cannot fully capture how contemporary technologies govern movement. To describe their temporal effects, this text proposes a ‘prediction of the now’. Rather than thinking in terms of velocity or acceleration, the prediction of the now upsets the linear timeline of data processing. The ‘now’ is predictively enacted through an anticipatory model of normality. With the prediction of the now, the subject under surveillance is no longer assessed in terms of its similarity to a ‘risky other’, or in terms of potential future harm; instead the other emerges by a logic of ‘surprise’ in the temporal variability between prediction and measurement. A hands-on examination of tracker software locates the temporal logic of surprise in the concrete functioning of the algorithms that produce a perception of movement. Specifically, the paper turns to the lines of code of DeepSORT, a multi-object tracker, which is tasked with describing movement out of immobile, frame-by-frame, detections. Drawing on the formulas that undergird these systems, this article explores the role of the surveilled subject in the governance of movement.
Introduction: The algorithmic perception of movement
Movement is central to the governance of places and people; it is both fundamental to and a threat to societies. In contemporary liberal societies, ease of movement is often seen as emblematic of individual freedom (Bigo, 2010). Whether at borders, or in city life, security professionals and their socio-technical devices determine who is and is not entitled to free movement (Bigo, 2014; Pallitro and Heyman, 2008: 319; Salter, 2013). Liberal security is thus a sorting practice, that produces ‘uneven mobilities’ (Gandy, 2021; Sheller, 2017). Control over movements is exerted not only at border sites (Amoore, 2006; Aradau, 2016), but extends into other realms of social life such as traffic flows (Coletta and Kitchin, 2017), city water (Latour and Hermant, 1998), urban crowds (Nishiyama, 2018), and building visitors (van der Maarel et al., 2026). Across these sites, technologies of sensing, interpretation and intervention qualify and shape movement.
Control over movement is thus harnessed to the ways in which technologies makes it governable. Over the past decade, a range of seemingly dispersed surveillance applications has come to rely on a particular kind of algorithm to ‘perceive’22 movement. Experiments with these systems were conducted, for example, to detect violent behaviour in the streets of The Hague and Rotterdam (Hulsen, 2024). At Berlin Sudkreuz train station the movements of people and objects were traced to assess whether luggage had been left unattended (Berliner Morgenpost, 2019). In-store foot traffic is analysed, warning managers of potential shoplifters (Luiten, 2024). Similar systems were to deter suspected burglars in a Rotterdam neighbourhood by changing the streetlights (Hamada, 2020). In another experiment, employees using a company parking lot were traced to alert staff of potential automobile thieves (Morellas et al., 2003). These cases represent a broader shift in security practices, in which camera based surveillance systems are augmented with algorithmic processing of video. Tracking bodies over time, these systems produce moving subjects.
Remarkably, many of these technologies are considered responses to earlier critiques of predictive and biometric surveillance by activists, civil society organisations and researchers. Institutions that research and implement these technologies claim that the shift to the capturing of movement is a watershed moment. They suggest that by refraining from identifying individuals based on static visual traits – such as faces and fingerprints – concerns about bias and privacy are no longer valid. Moreover, on a practical level, these technologies generally do not provide definitive assessments of their subjects, but assist security operators by filtering camera feeds, guiding their gaze and attention to those moments deemed to be of interest. This set of security technologies based on a subject’s movement would thereby constitute a safe, and ethically compliant technological advancement (e.g. Angelini et al., 2019).
But is it really? To assess this claim, we must know how the analysis of images over time enacts a moving subject. Moreover, if these systems indeed abstain from identification, how do they bring about the categorisations of suspicion that are to guide the operators gaze?
I suggest the use of such motion tracking technologies certainly complicates common grounds for critique on the use of algorithms in governance practices. What is at stake, I argue, is a misunderstanding of how movement – and by extension, the moving subject – is made calculable by algorithmic systems.
My argument has three implications for critical studies of security and data. The first implication is that we need to rethink suspicious categories based on the immediacy of observation, or the teleology of future harm that underlies predictive practices. To address the temporal effects of the algorithmic perception of movement, I propose the ‘prediction of the now’ through which the paper explores how the algorithmic perception of movement alters the security logic of anticipation, and augments it with one of surprise. The moving subject emerging from this process is no longer assessed in terms of its difference from others, but rather how it deviates from what it was expected to be. With the prediction of the now, the algorithmic perception of movement poses a distinct regime of security governance that renders the harmful ‘other’ in temporal terms.
Perhaps more fundamentally, the reconceptualisation of movement implies that socio-algorithmic systems of security cannot be reduced to a single temporal logic, we need a more fine-grained sociology of security practices and systems to derive, from empirical analysis, which specific logics might be at play. This means that grand narratives about algorithmic technologies should be re-assessed in light of detailed empirical observation of both the code and the social uses of such security systems. To that end, I turn to an algorithmic technique that plays a central role in many of these camera surveillance technologies: multi-object tracking. Multi-object tracking performs a seemingly simple task: given a sequence of images — for example video frames – it describes the movements of bodies. I engage with computer code to examine the operations by which snapshots coalesce into a mobile subject, and trace the conceptions of movement and suspicion they put to use.
Finally, drawing on the core tenets of signal processing that can be found in the inner workings of algorithmic movement tracking, the paper explores how understanding movement in light of a predicted present reconfigures critique of algorithmic security. It reframes the specific operations of the tracker to open up new avenues of complacency with and resistance to such surveillance practices.
Immediacy and anticipation
Central to this paper is the shifting logic in governance technologies that operate over time to examine moving subjects and attribute suspicion. The literature on the temporal effects of technology on security governance, has generated important insights about what might be at stake with such movement-based security systems, but they fail to account satisfactorily for the implications of the algorithmic perception of movement.
A first theorisation of the temporal effects of algorithmic governance is embedded in approaches which understand the monitoring of movement in light of the ‘real-time’ operation of algorithmic surveillance infrastructures. Such analyses describe the effects of observation in terms of the system’s velocity, or its derivative, acceleration. From this perspective, real-time monitoring technologies either allow for more speedy, efficient management of flows of people (Dijstelbloem et al., 2011) or conversely, that they fail to live up to claims of efficiency and instead add to the congestion in border control or police work (Leese, 2020; Leese and Pollozek, 2023). What these descriptions share is that surveillance technologies variegate the rhythms of social life (Coletta and Kitchin, 2017), and in their ever-accelerating operation, they exacerbate power while depoliticising decision-making (Møhl, 2020: 111; Olmstead, 2025). Using these devices, security professionals aspire to a form of ‘live governance’ by acting on events as they unfold, carving out situations in space and time deemed relevant for intervention (Walters, 2017). In this context, the acceleration of governance that these technologies make possible – decreasing the delay between sensor readings, their processing, and subsequent intervention – thus produces a sense of ‘liveness’ that affects the categorisation of suspicion.
In these accounts, movement is an effect that emerges from the immediacy by which the identification of subjects and events is accomplished (e.g. Kitchin and Dodge, 2011: 99). The accelerating pace at which such algorithmic technologies make surveillance data available rearticulates the relationship between security operator and its subject. ‘Rather than the periodic “stocktaking” of conventional statistics’, argue Isin and Ruppert (2020: 10), ‘populations are divided into clusters that are live and have pulses, flows and patterns.’ The ‘live’, moving subject is seemingly without a past or future, but emerges by virtue of ever-accelerating data capture and computation.
A second approach to the temporal effects of security technologies, has examined the pre-emptive techniques by which the likelihood of harm is calculated for subjects and situations. According to this view, algorithmic security technologies and thinking offer a probabilistic relation with the future23. In a sense, the prediction and pre-emption of crime (precrime) popularised by the 2002 film Minority Report, based on the 1956 novel by Philip K. Dick, serves as a reference point for the public and academic imaginary. This perspective undergirds practices such as predictive policing (Kaufmann et al., 2019) and risk profiling (Amoore and De Goede, 2008). Anticipatory action (Anderson, 2010) in the present is warranted based not on what an individual has done, but on the harm they are likely to inflict in a predicted future. As only limited data contains indicators for threats, pre-emption works speculatively (Amoore and De Goede, 2008; Massumi, 2007).
With these pre-emptive techniques a logic of anticipation structures time. The subject is place on a linear timeline, as data of the past is used to calculate risks scores that enable action in the present by speculating on a potential bleak and dangerous future (de Goede et al., 2014; Lyon and Wood, 2021; Massumi, 2010). Here, movement is not perceived as a sequence of snapshots, but emerges from a risk forecast that imbues the subject with potential for future harm.
While in both real-time and pre-emptive regimes of governance of security subjects are moving – they know temporal variation – through the back door their conceptualisations of movement sneak in sedentary notions of the security subject (see also Huysmans, 2022), that are accumulated and assessed on a linear timeline of past, present, and future24. While movement is primal to the temporal structures of both real-time and pre-emptive security regimes, both regimes share the ‘accumulative logic’ of databases. This is central to Haggerty and Ericson’s (2000) ‘data double’: an additional self produced through the aggregation of all kinds of digital traces, which make the individual hypervisible to surveillance (see also Raley, 2013a; Iveson and Maalsen, 2019). In practices that rely on this database logic, ‘flecks of identities’ sucha as travel logs, parking tickets and financial transactions are ‘drawn together’ (Fuller, 2007). As heterogeneous data is put in formation, an individual is assembled from an ever-growing collection of entries (Perret and Aradau, 2024), memos or notes (Bonelli and Ragazzi, 2014). Linking the individual and its records however are identifiers such as the face, fingerprints or a social security number, which are presumably immutable.
In addition, to assess these assembled, collaged, subjects, security technologies mobilise geometric distinctions: producing categorical distinctions among those who are likely to be ‘inside’ or ‘outside’ a defined group, territory, or cluster. Dissimilarities between subjects may be ‘banal’, as they can rely on ‘almost insignificant details—the time or length of a phone call, an overnight stay, or rare use of a mobile device.’ (Aradau and Blanke, 2022: 71) Such recorded traces, sedimented into distinct records, are mathematically rendered as singular ‘vector’ representations. In the process, the subject’s past movements depict a present state. Cast into these spatial representations, they can be compared by means of distance metrics. These comparisons yield a degree of ‘betweenness’ that describes the dissimilarity of a subject with others in spatial terms (Aradau and Blanke, 2017). Thus, while the subject is drawn together out of records that are collected over time, temporality vanishes in the geometric comparison between subjects.
In the technologies central to this paper, the subject’s movements require a computational step in which the immediacy of observation is complemented with a simulation of the present. An algorithmic perception of movement thus does not emerge by concatenating discontinuous data points at ever-increasing speed, and with ever higher resolution (see also Soon, 2019; Chun, 2008: 19). The ‘algorhythmic’ (Miyazaki, 2018) effects of the analysis of moving bodies invoke more convoluted timelines. Incorporating movement in security governance is more than merely an ‘expansion of the present’; time is structured on nonlinear timelines that disrupt relations among the historical archive, real-time monitoring and future-oriented prediction (Weltevrede et al., 2014: p181). ‘At issue,’ suggests Huysmans (2023), discussing contemporary security governance ‘is not that bodies are moving but that life exists only in movement.’ (Huysmans, 2023: 194) To gain purchase on the governance of movement, we thus need to imagine alternative understandings of its temporal effects that open up these sedentary security logics of accumulation and differentiation.
Security’s subject as a signal
The main claim of this paper is that the algorithmic governance of motion brings about a different logic of movement – and by extension, the moving subject – as a regime of security government. These technologies do not solely offer traditional ‘real-time’ observation, nor do they anticipate and pre-empt future catastrophe. Rather, the algorithmic perception of movement brings about a temporal regime of governance I term the ‘prediction of the now’.
The prediction of the now can be found in the interlocking timelines of observation and prediction, of immediacy and anticipation. The coming together of these temporalities produces discrepancies that quantify the subject’s attunement to what it was expected to be. The prediction of the now happens over and over again for each snapshot and for each video frame. In this cycle, this regime of governance prioritises the individual’s temporal variation, as it measures the ‘conformity of things to themselves’ (Huysmans, 2023: 197). Suspicion then, emerges when this conformity is not found.
Here, the suspicious subject lies not in the spatial distance – the ‘betweenness’ – of distinct subjects, but rather emerges spatiotemporally, at the onto-epistemic interstices of the predicted now, and measured data. This onto-epistemological gap has long been conceptualised by information theory and signal processing, fields that provide the frameworks and techniques on which the algorithmic governance of movement is seated. Inn these fields, the derivation of information from a signal may be thought of as the resolution of ambiguity or uncertainty that arises during interpretation. In other words, the moving subject is assessed for the predictability of its movements, quantified as the amount of ‘surprise’ they generate.
To comprehend how movement is governed, in Table 2, I distinguish the prediction of the now as spatiotemporal regime of suspicion from the spatial categorisations of suspicion that are to be found in real-time monitoring and pre-emptive accounts of governance. Note that in practice, these differences are not absolute. These regimes appear in entangled ways in complex governance assemblages. Yet, in what follows, I explore how a regime of the predicted present, with a temporal logic of surprise, can help us to consider the effects of algorithmic governance of movement. To that end, I turn to the case of the ‘Multi-object tracker’.
| Real-time governance | Pre-emptive governance | Prediction of the now | |
|---|---|---|---|
| Technology | bounded space (premises) | speculation | signals |
| Temporality | immediacy/near-past | predicted future | predicted present |
| Calculative technique | acceleration of data transfer | accumulation of data points | integration of state |
| Governing logic | observation | anticipation | surprise |
| Assessment | transgression | threat | conformity |
| Object | situation | spatial other | temporal other |
| Frictions | slowing down | exposing bias | widening the distribution |
Approximating continuity
Considering the state of technological development today, using algorithms to track the movements of people might seem deceptively simple. There is a wide off-the-shelf availability of object detectors, which locate objects like people, cars, or pizzas25 in digital images. Running an object detector for a series of consecutive video frames and putting these detections in sequence appears to produce a recording of a moving figure (see figure 9). Thus, as long as one would capture a person with sensors, producing a sequence of snapshots, movement seems to appear automatic. But does it?
What seems simple on the surface, is in fact a trick of our minds: persistence of vision (or flicker fusion) leads the mind to fill in the movement. The difficulty of this trick is exemplified by a canonical example from cinema: the wagon wheel effect. It is a visual glitch in which a spoked wheel might seem to rotate in reverse, depending on the angular distance between the spokes, the rotational speed of the wheel, and the camera frame rate (the time interval between recorded frames). The wagon wheel effect occurs when human vision confuses the spokes. Algorithmic object tracking faces a similar challenge: when multiple figures appear in the frame, or when one figure is occluded by another figure or object, like a lamppost, the figure’s movements become ambiguous. It becomes impossible to distinguish appearance from re-appearance across frames (see figure 10). Due to the sampling of reality into individual video frames, and quantising images into distinct pixels (converting a continuous range into a set discrete of values), the video stream is a discontinuous construct. Continuous motion therefore ‘becomes itself a secondary effect of discrete enumeration to be approximated’ (Ernst, 2011: 246). Discontinuous sequences of images or detections, appear as continuous ‘streams’ to the human eye due to the micro-temporal scales at which they operate (Soon, 2019). However, to make motion algorithmically, an additional computational step is necessary.
The algorithmic approximation that effectuates persistence of vision
is known as multi-object tracking. The multi-object tracker
takes as its input a sequence of video frames, for example from a
surveillance camera, and yields a set of trajectories that describes the
movement of visible bodies. It is used in algorithmic surveillance
technologies that examine movements, ranging from the basic task of
counting people going in and out of a space, by tracking the direction
of their movement, to systems that distinguish normal from anomalous
movements. ‘Multiple’ here denotes that several occurrences of a kind of
object can be in the frame at once. To produce trajectories, the
algorithm links the detected objects of a particular class
(e.g. person or vehicle) across the time gap
between consecutive frames.
The multi-object tracker can be considered an assemblage of algorithmic techniques, which are ‘at the same time material blocks of technicity, units of knowledge, vocabularies for expression in the medium of function, and constitutive elements of developers’ technical imaginaries.’ (Rieder, 2020: 16). The most canonical means of multi-object tracking is the Simple, Online, and Realtime Object Tracker (SORT) (Bewley et al., 2016). As an algorithmic technique, SORT circulates and mutates, and now has numerous implementations and derivatives. One variation is DeepSORT (Wojke et al., 2017), which augments SORT with an artificial neural network to compare detections. This architecture — SORT plus a neural network — again knows many derivatives, some of which are considered state-of-the-art when it comes to multi-object tracking (Papers with Code, 2024). In this paper, I turn to the reference implementation of DeepSORT that was released in parallel with its paper26 under the GNU General Public Licence v3.0. The open availability of the code makes it possible to apply the original implementation, written in Python, in my own codebase with which I produced the visualisations of the figures.
To analyse the multi-object tracker, the algorithmic technique that is at the core of the computational processing of movement, I embark on a technography of the code, a close examination and implementation of the underlying lines of code27. It is a form of exploratory programming (Montfort, 2016), that draws on software studies (Fuller, 2008), critical algorithm studies (Seaver, 2017), critical code studies (Marino, 2006) and media archaeology (Parikka, 2012). In what follows, explore the source code files to trace the lineages of the operations that turn distinct detections into subjects that are variable over time. This study thus mobilises technical description of the motion tracker as a means of teasing out how they configure, and are configured in, security governance.
Moving bodies in four lines
To start working with the DeepSORT tracker, one only needs a snippet of code: 28
for frame in video:
detections = YOLO(frame)
tracker.predict()
tracker.update(detections)Running these four lines of code sets off a cascade of operations.
The first line, a construct known as a for loop, describes
an iteration over a video, in which each frame is considered as a
distinct snapshot29. On line 2, the frames are passed
to an object detector — the commonly used Ultralytics YOLOv8 (Jocher et
al., 2023). Reading from right to left, the result of the
call to YOLO is assigned to the detections
variable. The result describes a list of image coordinates (x and y
values) that represent the top left and bottom right corners of a
rectangular bounding box that locates the detected object in the image,
as well as a category – defining the detected object as a
person, a car, a pizza, or
another of the 80 categories of the COCO dataset (on categorisation, see Malevé,
2020). Lines 3 and 4 run the DeepSORT object tracker on the
detected bounding boxes. It is these two lines —
tracker.predict() and tracker.update() — that
effectuate algorithmic persistence of vision, transforming static
snapshots into movement.
Line 3: predict() the now
The object tracker describes the trajectories of those who passed by
the camera, yet paradoxically, even before the detections are presented
to it, it invokes the forward-looking logic of prediction. Those
familiar with coding might wonder why the tracker needs two function
calls. Why isn’t a single tracker.update(detections)
function sufficient? Examining the operations that come into play in
calling the tracker’s predict() function can reveal the
entangled temporalities in the monitoring of movement. This means we
need to step away from self-written code, to the code as written in the
DeepSORT implementation (lst. 1).
Listing 1: Excerpt from DeepSORT. Running predict() on
the tracker object, calls the nested predict() function on
the tracks stored in the tracker’s state.
In this first snippet, the discontinuity of video is immediately
apparent. The two lines of documentation that are part of lst. 1 (indicated by """),
describe when the predict function should be used (‘before
update’) and stress how time is punctured as it is made to
pass in a sequence of steps. This resonates with what Deleuze (2024), following Bergson, discussed on
cinema: the equidistant perforations on the film-strip produced a
uniform time interval between frames. Each frame of the video is an
‘immobile [section] of time’ from which movement is to be reconstituted
(Deleuze,
2024). The tracker invokes yet another, nested,
predict() function for all the tracks that it has at that
moment in memory (the self.tracks variable refers to data
kept in memory by the tracker; its present state). This suggests some
information is carried over from the previous time step to the current,
for which the predict() call is executed. Thus, to overcome
discontinuity and stillness, the tracker has to ‘propagate’ previous
measurements into the now.
Listing 2: Excerpt from DeepSORT. A track’s predict()
function does not yield an absolute position and velocity, but, by
representing these values as probability distribution, provides an
assessment of the uncertainty of the prediction.
The origin of the tracker’s predictive logic becomes apparent when
tracing the code invoked by track.predict() (lst. 1 line 55). The function’s definition
can be found in another source file (see 2). The comments in the code refer to the
Kalman filter. This technique, devised for the Apollo mission (Kalman,
1960), is one of the most used signal processing techniques
in control theory to process discrete time-series data (Musoff and Zarchan, 2009). It is used
not only in camera security, but also ‘to control a vast array of
consumer, health, commercial, and defense products.’ (Grewal and Andrews, 2010). It integrates
multiple signals from GPS satellites into a single location and
velocity. It can estimate the movement of objects in the interval
between individual radar detections. Missile defence systems use it to
anticipate a projectile’s trajectory. In the case of the movement
tracker, the distinct measurements that come from the object detector
represent the position of a body. They do not contain the essential data
on its velocity or acceleration that describe a body’s movement. The
Kalman filter can mathematically infer these properties by subsuming the
measurements of position into a predictive model of reality. The
multi-object tracker overcomes discontinuous time by constructing the
present via a simulation that ‘stretches forward’ observations across
the temporal gap between frames. In other words, past tracker data is
used to predict the now.
The Kalman filter’s predictive process is developed to operate in an inherently noisy system. The filter thereby operationalises Shannon’s mathematical theory of information, which states that any input value is distorted by aspects of reality one does not want to measure (Shannon, 1948). In case of the object tracker, the readings of an image sensor contain electrical noise, the compression of the image pixels by means of a video codec alters their values, and many more such distortions are bound to happen for each processing step. Likewise, the detection algorithms that locate the body in the still frame are assumed to be imprecise as they are, for example, impacted by occlusions. In a noisy system, whether data comes from an object detector, or any other sensor, no measurement can be considered absolutely certain. Therefore, while reality can be measured, due to the noise in the process, it is presumably ‘hidden’ from observation.30 To come to terms with these uncertainties, the logic of measurement is shifted. In this context, although reality is hidden or unobservable, it is ‘latent’. In medicine, the latent period of an infection is the time exposure and symptoms, or between drug intake and effect (Veel, 2021). Similarly, a latent reality and its movements are dormant: they are still undeveloped, but bound to develop. The assumption then is that these innate developments of the latent subject can be modelled.
If we thus understand the predictive process in light of these core
tenets of signal processing, the Kalman filter filters the measurements
for ‘relevant’ data – the signal – by quantifying the noise in the
system. Therefore, when the Kalman filter predicts the body’s position,
and its first derivative over time – velocity – it does so as a
multivariate probability distribution. The prediction covers a range of
likely next positions centred on the mean, the
spread of which depends on the modelled certainty, the
covariance. By simulating the body’s most likely movements,
the temporal logic of the forecast is invoked to anticipate the most
likely next step, simultaneously quantifying the confidence of that
prediction.
What is the role of this prediction in the broader functioning of the
tracker? As the earlier snippet of code (lst. 1) shows, the predict()
function does not return its predicted values for use by the end-user of
the tracker (i.e. the programmer). Rather, the predicted present stays
within the logical confines of the multi-object tracker’s tracks. This
implies the forecast – the anticipated present – needs to be understood
in the recurrent operation of the tracker, in which the same steps
repeat for one frame after the other. In this operation, which runs for
sheer infinite times when seen from the micro-temporal scale at which it
takes place, the predicted present serves as a predisposition
that shapes how new measurements are incorporated in an additional
computational step.
Line 4: update() and the recurrent integration of
state
The function of the predicted present thus becomes apparent in the
subsequent step of the tracker, as it is fed the detection data of the
latest frame. Here, the predicted present provides the ground upon which
this incoming data is assessed; the result of this assessment becomes
the input for the next prediction. To that end, the
update(detections) of the multi-object tracker performs two
main tasks. A call to match() associates the incoming data
with existing tracks. After associating, a call is made to the second
part of that track’s Kalman filter in another, nested,
update() function. The Kalman filter’s update integrates
the associated data into the track state (see lst. 3), by which the next timestep’s prediction will
be made. It is in the recurrent integration of state that a
different formulation of movement becomes apparent; one that is anchored
in the subject’s temporal variability.
Listing 3: Excerpt from DeepSORT. The update() function,
in the Tracker object. First, the _match()
function assess whether any of the detections are similar to the
predicted positions of known tracks. In a second step, matched tracks
are updated with the associated detection by calling the Kalman filter’s
update step.
# Run matching cascade.
matches, unmatched_tracks, unmatched_detections = \
self._match(detections)
# Update track set.
for track_idx, detection_idx in matches:
self.tracks[track_idx].update(
self.kf, detections[detection_idx])
…In match() (see lst. 3 line
98), the tracker associates new measurement data
(detections) with previously processed tracks based on
their movements. With the original SORT technique, the predicted present
positions are compared with the newly detected positions based on a
distance metric. The detections are associated depending on which track
is nearest in the two-dimensional image space (the Hungarian algorithm).
If a position cannot be associated, a new track instance is created; if
a track has no matching detection, it is marked ‘lost’. With DeepSORT, a
feature vector augments the two-dimensional distance metric. This vector
is a numeric description of a small portion of the video frame, obtained
by cropping it to the detected bounding box and feeding that to a (deep)
artificial neural network31. The similarity of
feature vectors is again determined using a distance metric. However,
the vectors involved in the comparison arise from distinct temporal
logics: a measured now, and a predicted now, which was modelled after
past measurements. As the predicted now is not merely a point, but a
distribution, the distance metric is not purely geometric but reconciles
the statistical likelihood of the measured position. Moreover, to
associate data with subjects the tracker thus does not rely on a body’s
static identifiers, such as fingerprints, or faces. Rather, the value of
this feature is updated on each iteration of the tracker. In the
recurrent matching, the identifying features of the moving body change
as the movement unfolds.
The subsequent Kalman update() function recurrently
integrates measurement data into the tracker’s state (see lst. 4). The quantisation of movement into distinct
frames, and the noise inherent in these observations, poses a
computational challenge: any prediction of a body’s position based
solely on the last measurement is considered highly unreliable. However,
the Kalman filter cannot simply accumulate all previous data points to
run its calculations. Multi-object trackers are made to run on ‘edge’
devices (rather than central cloud servers), such as cameras or embedded
computers, thus both memory and computing power are limited. Retaining
the full history of an object for the prediction of their next position
would both require an increasing amount of memory and exponentially
lengthen the amount of time needed for the prediction. This increase in
computational complexity impedes the low delays upon which the tracker’s
‘nowness’ relies. To overcome these limits, the process of integration
compresses the sequence of observations, each of which merely captures a
body’s still position, into a behavioural model of the moving subject
(cf. Pasquinelli and Joler, 2020a). As this
state should not represent how things are, but what they most
likely will become in future timesteps, the Kalman filter augments the
measurements with the derived velocity. Like the latent spaces in more
complex generative neural networks, the filter’s integrated state
fulfils ‘a seductive political promise of locating the hidden tendencies
in populations, places or scenes’ (Amoore et
al., 2024: 8). Thereby, as data comes in, the integrated
state shapes incoming data; and, as the state is used to predict the
subsequent timestep, it is itself shaped by that data.
Listing 4: The update() function of the Kalman filter,
which calibrates the filter state.
projected_mean, projected_cov = self.project(mean, covariance)
chol_factor, lower = scipy.linalg.cho_factor(
projected_cov, lower=True, check_finite=False
)
kalman_gain = scipy.linalg.cho_solve(
(chol_factor, lower),
np.dot(covariance, self._update_mat.T).T,
check_finite=False,
).T
innovation = measurement - projected_mean
new_mean = mean + np.dot(innovation, kalman_gain.T)
new_covariance = covariance - np.linalg.multi_dot(
(kalman_gain, projected_cov, kalman_gain.T)
)
return new_mean, new_covarianceWhere in the predict() step the confidence of the
prediction is assessed, the update() function – which in
the recurrent operation both follows and precedes predict()
– is to quantify the degree to which the body’s movements adhere to its
derived model. As outlined in the previous section, the information
theories that underlie the Kalman filter assume noise to be present in
the measurement. Moreover, the Kalman filter’s model, compressing all
observations into a single representation, is assumed to be lacking.
Consequently, the approximated movements of the body are never exact. A
metric central to information theory, surprise (or,
‘information content’, denoted as innovation in the code,
lst. 4 line 192) describes this difference
between the predicted present and the observation data. The Kalman
filter’s update() step uses this difference, the surprise
yielded by the coming together of measurement and prediction, to adapt
the weighing (the kalman_gain) of the simulated positions
versus the latest measurement data (Musoff and Zarchan, 2009). By
recurrently tweaking the mean and covariance, the filter minimises the
expected surprise of future measurements by minimising the statistical
difference between what was expected and what was observed. Measuring
conformity of the movements to its model, surprise
operationalises the gap between the heterogeneous temporal logics of the
predicted present and the immediate observation.
After the integration of state, the cycle resumes. All that is
carried over to the next iteration of predict() and
update() is the inferred present state and its uncertainty.
In this recurrent process, gathered knowledge is assessed by its virtue
to predict subsequent positions, while the subject is assessed by its
adherence to the model. In the prediction of the now, the moving
subject, temporal variable and fundamentally unknowable, is governed as
a signal.
Conclusion
This paper engaged with the governance of movement, and how it invokes recurrent processes of prediction of the now and the integration of state. These operations suggest two main differences between the temporally variable signal subject and the sedentary subjects of real-time and pre-emptive regimes. The first concerns the role of prediction. The predictive techniques employed in pre-emptive security technologies simulate harmful futures. Assessing these predicted futures and quantifying their harm establishes a regime of anticipation that warrants action in the present. We should, however, avoid conflating predictive algorithms with such a ‘future politics’ (Amoore, 2013). While the ‘prediction of the now’ uses similar simulation techniques, it differs both in its objective and in its temporality. The prediction of the now comes with a temporal regime of anticipation that is no longer oriented at the future, nor does it concern the marginally likely, catastrophic futures. Rather, it uses signal processing techniques to overcome the onto-epistemological challenges of information theory, in which there is no way to consolidate the latent present without prediction. Anticipation reconciles the continuously changing subject and makes it amenable to governance.
The second concerns assessment and suspicion. Much like under real-time and pre-emptive regimes of governance, the integration of state calculates the similarity between subjects in a ‘geometric’ feature space. However, the entries in this comparison differ. In the recurrent operation of the tracker, the modelled individual – including its measurement and prediction uncertainties – is calibrated such that it is best able to describe the next measurement. Hereby, as the Kalman filter minimises the expected surprise, it assesses how measurements of the unknowable, latent, moving subject adhere to their predicted present. Hence, it does not differentiate who is part of which group, or draw categorical distinctions between subjects and spatially distinct ‘risky’ others. Instead, with the prediction of the now, the subject is assessed for ‘conformity to oneself’ by contrasting it with its temporal other.
Subjects are thus governed as signals, shifting techniques of observation into prediction, and of differentiation into integration. What the logic of surprise makes apparent is that ‘algorhythmic’ systems, more than governing the velocity of processes of movement (Leese and Pollozek, 2023), anticipate these rhythms. The failure to predict the off-beat move produces a sensory impulse. Surprise is visceral; it catches the eyes and ears of the observer and raises their eyebrows. Surprise sharpens attention.
Thinking in terms of the prediction of the now, and the governing logic of surprise, has consequences for the critique of algorithmic security practices. For example, algorithms used in security inevitably err and the failure to predict has provided an important locus of critical inquiry (see van de Ven et al., 2026). From the perspective of real-time security, in which algorithms tend to be used for identification of subjects or situations, the failure to predict leads to a rather unambiguous error. That is to say, a recognition system either recognises the subject or it does not. Even though in practice it might not be black-and-white – a facial-recognition system can provide a list of likely subjects instead of just one – the outcomes of such a system can be assessed in terms of its precision and accuracy. Critique then, highlights the imbalance of error rates between different populations: algorithmic bias. In pre-emptive security, the role of errors is much more ambivalent; if a forecasted disaster occurs, it proves the risk model right, if it does not occur, it proves the pre-emptive intervention successful (Calhoun, 2023). While the inevitable failure to predict in pre-emptive security practices evades critique, it nevertheless is an unintentional and undesirable side effect.
With the prediction of the now, however, the error inherent to statistical prediction is granted yet another status. As the subject’s steps are subordinated to a model of their likeliness, deviating from the prediction is no longer a failure of the technology. With surprise, the discrepancy between prediction and measurement becomes a necessary characteristic of the system’s operation that describes the subject’s failure to ‘conform to oneself’.
As much as the logic of surprise evades critique, it also makes alternative loci of critique visible. With the prediction of the now, a crucial role is granted to the generative model of normality that produces the temporal other. Tis model of normality is constituted by enfolding each tracked step of the moving subject. Consequently, in the prediction of the now, those observed by tracking devices collectively produce the harmony that is to be secured. This account of the prediction of the now thereby resonates with analyses that underscore that there is no security without insecurity, and that it does not suffice to focus analytical attention solely on the power of those observing over those observed (e.g. Bigo, 2008). Many of these analyses, however, develop preventive accounts of security, in which liberal democratic principles of innocence until proven guilty, are replaced with an a priori suspicion (Bigo, 2025), that ‘enlist[s] everybody under the category of suspicion’ (Aradau and Van Munster, 2007: 34), turning them into persons of ‘national security interest’ (Amoore, 2013). In contrast, the prediction of the now reveals a more reciprocal relation: the conforming subject helps constitute the norms against which deviations are measured. Every surveilled step, is an ‘act’, however little, in a diffuse securitising process (Huysmans, 2011; on walking as act, see Certeau, 1984), that reconfirms, or slightly adjusts the model by which the present is anticipated. Seen through a lens of surprise, being watched is not the same as being noticed by surveillance. As anyone’s patterns of movement become the backdrop against which those moving unpredictably stand out, conforming to predictability becomes indistinguishable from complacency with surveillance.
At the close of this analysis, therefore, I would like to deliberate on the implications of a logic of surprise by considering what might elude attention. To that end, I turn to the equations that undergird the algorithmic perception of movement.
The Kalman filter, employed by the tracker, predicts a subject’s present as a probability distribution. As such, it assesses the subject’s predictability by calculating the spread of the distribution of the predicted properties. This means that when uncertainty increases, the spread of the prediction widens. Therefore, perhaps paradoxically, the deeply unpredictable subject will yield a lower surprise metric, as long as it conforms to its unpredictability.
As such, in terms of eluding attention, we can draw inspiration from two figures that inhabit the edges of predictability. Going against the linearity of time and movement in digitisation processes, Roos Hopman (2024) turns to snails and their remarkable ability to drift. Drifting implies renouncing a meaningful course; embracing instead an aimless movement. Unlike the loiterer, a rather motionless figure, the drifter is constantly moving. In that sense, the movements of the drifter can be related to those of other wandering figures, such as the flâneur, the tramp, and the vagabond, all of whom have historically been unwanted in public space (Buck-Morss, 1986). By contrast, where the drifter ends up in unexpected places by passively following the current, Shintaro Miyazaki (2023) suggests a hyperactive mode of ‘loosening up structures’ (2023: 10) by engaging in counter dancing. Dancing, Miyazaki suggests, is a social practice, that moves in unexplored directions. Both the dancer and the drifter, in either hyperactivity or passivity, lack directionality. Their trajectories maximise the uncertainty-term in the predictive equations of the tracker, thereby minimising the degree of surprise.
Ultimately, embracing the drifter’s wanderings and the dancer’s spontaneous exploration, offers a way to disrupt the recurrent loop of surprise. These modes of existence, outside the constraints of expected behaviour, intervene in the reciprocal relationships between individual and population, and between observer and observed. Resisting the algorithmic perception of movement lies neither in slowing down the system, nor in evading recognition, but in accepting the unpredictable, and maximising the uncertainty it seeks to minimise.
5 Perplexity
Surveilling through indifference
This text complements the interactive installation Perplexity. Documentation on the installation can be found in Annex A. The thinking presented here was developed through the hands-on work of programming and problem-solving that was involved in making the installation. It has been published in abbreviated form in Peer-reviewed newspaper: Everything is a matter of distance (2025), 14(1), 29–30. ISSN 2245-7607. PDF and Online
Cameras have become ubiquitous in public space. In city centres, shopping malls, or train stations, camera surveillance sets out to spot “deviant” behaviours to detect or pre-empt unwanted events. However, the increasing number of cameras produces so much footage that there are not enough eyes to constantly monitor all video feeds. Often, one person is responsible for more than a hundred simultaneous streams. The past years have seen the introduction of algorithmic techniques in observer rooms that are to guide the operator’s eyes by singling out particular behaviours. A detailed examination of the operations by which these – sometimes speculative – security technologies make behaviours in public space visible to an operator can open up ways to think about the relationship between surveiller and surveilled that reach beyond the minutae being discussed.
Many analyses into the effects of both camera surveillance and algorithmic surveillance can be traced to philosopher Michel Foucault, and his work on the panopticon. The panopticon is a prison designed by Jeremy Bentham in 1791. It is a circular building without doors, in the centre of which is a watchtower, from which a prison guard can see into all cells. Foucault used the panopticon as an analytical device to describe the normalising effects of disciplinary power that works through a central observer (Foucault, 1977). Most crucially, Foucault made palpable that it is not relevant whether an observer is actually present to monitor its subject; the mere idea of being observed is enough to keep people in check. Due to the hierarchies between surveiller and surveilled, the subjects of surveillance internalise the vision of the other. Foucault had not intended to develop a normative account of disciplinary power – he did not consider this kind of power inherently good or bad – instead focussing on the mechanics by which this kind of power subtly develops as a diffuse network of people and techniques. Nevertheless, his work has often been used to describe the oppressive effects of surveillance. From this perspective, surveillance is characterised by an authoritative observer, and observed subjects can but bow their heads and walk the line. The Foucauldian-inspired critical narrative of surveillance, anchored in this binary between surveiller and surveilled, has trickled into a general imaginary of surveillance as an invasion of privacy, and a limitation to freedom of behaviour enacted upon the individual by, primarily, the state.
At first glance, Foucault’s description of the central observer in the panoptic prison, maps neatly onto contemporary camera surveillance: all video streams flow to a central hub, in which a camera operator might or might not be looking at these images. While the archetype of the panopticon provides a very useful description of the normalising effects of many security and surveillance practices (Galič et al., 2017: 18), such a one-to-one mapping would be too simple, and fails to account for all features of CCTV surveillance (see e.g. Lianos, 2003; Davidshofer et al., 2017). For example, in his ‘Postscript on the Societies of Control’, Deleuze highlights in contemporary society, bodies are not solely moulded within confined spaces such as the prison, the school or the military camp – as they are with disciplinary societies – but control can just as well take place at a distance, in open spaces, by regulating access (Deleuze, 1992). Drawing on Deleuze, Haggerty and Ericson similarly highlight how contemporary surveillance relies less on a single, central observer, but instead a ‘surveillant assemblage’ branches rhizomatically into all kinds of information sources, effectively marking the ‘disappearance of disappearance’: there is no escaping surveillance monitoring (Haggerty and Ericson, 2000). These forms of surveillance challenge the top-down hierarchies of surveillance, in which the few control the many, opening them up for bottom-up surveillance and non-state security actors. Despite these crucial reformulations of surveillance, what persists is the normative surveiller-surveilled binary, in which those watching exert power over the watched.
When discussing surveillance in the public space, the self-disciplining implied by the panopticon and subsequent diagrams of surveillance, seems to tell only a partial story, for ultimately, most people simply do not care. The vast majority cares as little about being watched by the state as they care about the data gathering by ad companies. When one would ask a passer-by about camera surveillance, they might respond with surprise, or voice some obligatory comments of concern, but it’s only seldom heartfelt. They go about and do their business. Already in his 1984 description of “Wandersmänner” (wayfarers) in public space, Michel de Certeau (1984) suggests the “chorus of idle footsteps” traversing the city is largely indifferent to any top-down interference. Even I, a researcher of algorithmic security, shrug about cameras when I routinely cross the train station, only worrying about catching the next train home.
Algorithmic anticipation
The indifference exhibited by most of the surveilled suggests that framing them merely as ‘victims’ tells only a partial story. Rather, by examining how contemporary surveillance technologies negotiate deviancy and normality, I propose a reconfiguration of the subject under surveillance.
In surveillance practices, the notions of “deviancy” and “anomaly” serve as a catch-all category for any unexpected behaviour. Spotting such behaviour is often considered an art — a “gut feeling” conditioned by experience; or a sharp eye that some have while others don’t (Amicelle and Grondin, 2021; Norris and Armstrong, 2010). The threat models that warrant camera surveillance suggest the public needs to be secured from threats of terrorism or ‘high-impact crimes’. Training is thus tailored to recognise such specific threats by their distinctive ‘modus operandi’. However, in their everyday work, rather than mobilising concrete future scenarios, security practitioners rely on anticipation, as they relate “almost in a bodily, physical manner with ‘risky’ and ‘at risk’ groups” (Bonelli and Ragazzi, 2014) to mark people as ‘out of place’.
With the introduction of algorithmic deviancy scoring, the construction of anticipation needs to be reconsidered. Where a traditional machine learning detector is trained by example, such a setup struggles to deduce anomalous patterns from data for two reasons. This is, first, because the anomaly serves as a catch-all term for any nonconforming observation, and thus forms an open set of immensely heterogenous behaviours. For example, “in video surveillance, the abnormal events robbery, traffic accidents and burglary are visually highly different,” (Pang et al., 2021: 38:3), which makes it difficult for a mathematical model to converge. Second, due to the relatively rare nature of incidents, data available for training contains many more examples of “normal” than of “anomalous” behaviours. For example, CCTV footage of a shopping street will mostly show people calmly going about their business; moments that contain fighting or robberies will be rare. Traditional detection algorithms struggle with this class imbalance: these techniques perform best when each category (‘normal’ vs ‘anomalous’) contains roughly the same number of samples. Lastly, because any statistical outlier is potentially an anomaly, it is difficult to distinguish anomalies from mere measurement noise, assumed to be present in any analytical instrument (Chandola et al., 2009: 15:3).
To overcome these challenges, a logical reversal is invoked. Instead of detecting deviancy, normality is measured. Trained on vast quantities of “normal” data, a generative model uses past measurements to simulate the present. These forecasts (or, nowcasts) are then used to assess the likelihood of the movements observed. The anomalous is thus no longer considered in terms of proximity to a predefined ‘risky’ other, but as a measured distance from a simulated normality.
This unpredictability score resembles a metric known as perplexity. Perplexity, a concept from information theory, was originally introduced in the context of speech recognition, and now has become a prominent error metric for assessing algorithmically generated sequences – in particular of large language models. With perplexity, for each ‘token’ in a series — whether it is a word in a sentence, or a step in a trajectory — the predictability of that token in relation to what came before is calculated. As a measure of surprise, perplexity is the logical inversion of algorithmic anticipation.
With perplexity the present is governed through simulation. This simulation forfeits any relationship with a predefined ‘risky’ other, but rather defines it through a degree of predictability, that is, it represents normality. The jarring sensation of perplexity, of the failure to anticipate, is as visceral as the ominous sense of threat, operationalised by risk technologies (Massumi, 2010).
The erroneous forecast is no longer a bug that needs to be solved, but has become a feature. By subduing human steps to a model of their likeliness, it is no longer the algorithm that errs but the human who is deemed unpredictable. For de Certeau (1984) the trajectories of Wandersmänner elude legibility. With perplexity, it is precisely the lack of legibility that becomes an indicator for suspicion.
Routines
In our day-to-day routines, we travel set paths through streets, train stations and parks. Through perplexity, surveillance capitalises on these movements. They provide the training data upon which normality is forecast. Thereby we co-produce the backdrop of normality against which anomalous movement stands out (see also Pasquinelli, 2015; Canguilhem, 1978). There is no outside to surveillance. In the production of perplexity, everyone is implied.
The limits of the panopticon as a model for surveillance in public space thus become visible. Bentham’s architecture, and the subsequent analyses by Foucault and many that followed, rely on a clear demarcation between those in the tower and those in prison cells. These boundaries have now blurred. Rethinking the relation between normalcy and deviancy makes apparent that while everyone is watched by surveillance, the majority is not targeted.
Thinking through the logic of perplexity and integration of minor deviancies, however, reframes participation from a fragile balance to an actual benefit. A general public chooses convenience – not some sort of ‘false’ convenience with a hidden risk, but actual convenience – and thereby actively contributes to the background of normality against which the other stands out. Thus, critique of surveillance should refrain from (only) convincing people they are being harmed. Rather, they are complicit in constituting normality. Consequently, conformity to normality is complacency to surveillance.
As a final note, let us turn to Michel de Certeau, who conceptualised that, as much as those up-high in the skyscraper looking down – it is the movements of pedestrians that ‘write’ the city (Certeau, 1984). The pedestrian, the city dweller, the observed, shapes what is being seen and how that is understood. Thereby de Certeau liberates the walking subject: the pedestrian is no longer stifled by observation, but is granted agency in how one goes about the city. Walking not only affirms and respects, it can also try out and transgress. In the reciprocal relationship between individual and population, transgressions — breaks with predictability — are only momentary interruptions, to be enrolled in next forecast of normality. Collectively, we can make normality more unpredictable.
Conclusion
This dissertation started by asking: what is it that makes someone “stand out” to an observer? It examined the techniques by which algorithmic surveillance analyses movements to filter deviant behaviour. Along the way, it mobilised concepts from signal processing on which these techniques heavily draw, to understand the political effects of algorithmically surveilling movement.
The first section of the dissertation considered what happens in the onto-epistemological move by which data-gathering and processing are subsumed under a sensing practice. It problematised how such a gesture black-boxes the intricate cascade of operations at play in surveillance assemblages. A computational sensor produces its signal by by means of a sequence of transductions. Each transductive step of the processing pipeline selectively throws away input data, and recombines what remains. My co-authors and I argued these minutea matter for the politics of such a system. Black-boxing transductions depoliticises measured phenomena as if they are mere ‘emergent properties’, while each re-articulates the relationship between observer and reality.
The second section explored how social sciences can unwrap black-boxed security technologies. The section mobilised diagramming as a mapping approach. Diagramming complements interviews with practitioners who work with camera-based algorithmic security systems with drawing. By means of the resulting diagrams, the paper examined where and how these practitioners draw the relations of their devices. My co-authors and I highlighted how the predictions of algorithmic ‘vision’ are ever uncertain, thereby configuring their use in security practices around the ever-present risk of error. Moreover, the papers suggests algorithmically accounting for movement is not just the result of an increase in processing speed of frame-by-frame detections over time, but bears more complex timelines.
What then, is movement in algorithmic terms? This question is central to the third paper, which developed a technographic account of an algorithmic movement tracker. Movement tracking, the section argued, calls attention to temporal regimes of security and threat perception, different from the dominant accounts so far developed in security literature on computer technologies. Indeed, the ‘prediction of the now’ on which the algorithmic perception of movement relies, cannot be accounted for in terms of the immediacy of ‘real-time’ surveillance – which focusses on the velocity of data flows – nor the anticipation of future harm that is central to pre-emptive, risk based security. The prediction of the now interlocks logics of anticipation of normalcy with one of surprise. These logics can be traced to practices of signal processing, in which their conjunction resolves fundamental uncertainties of measuring reality by filtering information from noise. When tracked, the moving subject is treated as a signal, recurrently assessed for the degree of surprise its actions yield.
The last section of this dissertation explores the signal subject as a reconceptualisation of the relationship between the surveiller and surveilled. This exploration takes place across an interactive art installation, and an accompanying essay. The installation Perplexity, a re-modulation of an algorithmic anomaly detector, intervenes in passersby’s everyday routines. By subduing their set paths to algorithmic prediction it invites them to reflect on how the convenience and efficiency of their routines contributes to their surveillance. For how is one to act, and how can one act, in the face of complacency? The playful context of an interactive installation opens this question up for experimentation. The essay, in turn, develops the insights gathered in making Perplexity. Together, these parts combine personal reflection, technical description, and political analysis to come to an affective account of what it means to be observed.
Thinking through surveillance in terms of signal processing – of anticipation and surprise – the pedestrian is not stifled by observation, but granted agency. How one goes about the city alters how other’s movements are understood. Walking thus is a communicative act: as the passerby ‘gambols, goes on all fours, dances, and walks about, with alight or heavy step’ (Certeau, 1984) they rewrite the bounds of normality.
The signal subject
The anticipation of normality as a central logic to algorithmic camera surveillance, has consequences for how we, in both critical security studies and software studies, can think of targeting, and non-targeting, practices in security and surveillance. Critical literature on security, in particular on computational surveillance – dataveillance (Raley, 2013b) – and risk technologies, has highlighted how the surveilled subject is established as target by means of ‘differential normalities’. Under these differential normalities, similar to the prediction of the now, the subject is not assessed against some kind of general, disciplinary, norm. Rather, the suspicious subject is assembled by attributing signification to certain ‘footsteps’ that seem inconspicuous, or banal, when those would have been observed in isolation (Amoore and De Goede, 2008: 178; drawing on Foucault, 2009: 91). In their influential account on risk technologies, Amoore and de Goede highlight how such differential normalities make it possible to pre-empt potential futures, of ‘what could be terrorist schemes or attacks’. This leads them to conclude that even the ‘non-decision’, the decision not to follow up, not to intervene, is merely temporary: ‘The tracking and tracing of transactions of many kinds appears to allow for the perennial deferral of decision, for the always possible intervention in the face of imminent surprise or threat.’ (Amoore and De Goede, 2008: 180; Bigo, 2014: 219 draws a similar conclusion) In such an understanding of surveillance, ‘the act of targeting is an act of violence even before any shot is fired’ (Weber, 2005; in Amoore and De Goede, 2008). In these narratives of targeting, non-interference is burdened with permanent suspicion, placing the subject ever at risk of intervention.
Examining the logics of prediction and simulation that produce moving surveillance targets, I draw another conclusion regarding non-decision. To do so, I suggest, we have to consider a fundamental aspect of normalcy: its production. The production of normalcy becomes explicit when surveillance is automated. In the algorithmic analysis of human movement we find the immediate disassembling of deviancy, and its integration into a predictive model of normality held by the observer. In such an assemblage, the subject’s inconspicuous movements no longer feature as signifiers for potential future intentions. Instead, as the subject’s movements are folded into the surveiller’s predictive model they contribute to a baseline for normalcy. Herewith, the ever looming threat of misclassification does not fully disappear, but it no longer is a mere coin toss away either.
Consequently, we should thus not be surprised that most passerby are rather indifferent to their own surveillance, as the closing essay of this dissertation postulates. The rollout of camera surveillance continues at unprecedented scales: as much in cities, as at every other front door, where a cloud-connected Nest camera – or the like – observes anyone passing by the house. These developments are hard to account for with the hierarchies implicitly present in common descriptions of algorithmic surveillance, in which the subject is disciplined through observation, and ever at risk of misclassification. In such a rendering, security is intimately connected with control over the subject: “You are controlled for your own protection, you are protected so that you can be controlled.” (Gros, 2019: 180) However, for the gleeful majority the control exercised, and risks imposed, by camera surveillance – algorithmic or not – is virtually none.
This conclusion should not be confused with a normative statement of ‘good’ or ‘fair’. To be sure, the convenience chosen by a general public is not some sort of ‘false’ fragile convenience with a hidden risk, but actual convenience. Yet, it would be far too simple to equate the convenience of the majority with fairness. Nor is my claim that surveillance comes without a power imbalance; the problems highlighted by surveillance literature, in particular around the position of minorities, stands true and tall: for some groups, the convenience indeed is fragile, if it is experienced at all.
My point is that, perpetuating the implicit narratives of surveillance hierarchies, in which the observer exerts power over the observed, fosters a sense of powerlessness on behalf of the observed, and with that a sense of unaccountability. From such a perspective, when subjects willingly submit themselves to data based surveillance techniques, the experienced convenience unwittingly comes at the price of their own freedom (Bigo, 2014: 219), and thus they are as much the victim as the other. However, as subjects move through public space, I argue, we should examine the power relations beyond a mere surveiller-surveilled binary. Attention to the production of a model of normality, shifts the binary, top-down narrative of surveillance toward a ‘government of self and others’ (Gabriels and Coeckelbergh, 2019), similar to citizens’ voluntary participation in surveillance as they share their data on online platforms (see also Wood and Monahan, 2019). In algorithmic camera surveillance, under a regime of surprise, the subject that by and large conforms to the predictive model, plays an active role in the surveillance of others, exchanging the freedom of others for their own convenience. Conformity to normality is complacency to surveillance.
As this dissertation draws to a close, I somewhat speculatively propose an analysis of the algorithmic assessment of movements can help us rethink the constitutive role of the surveilled subject in more general terms32. Early critique on algorithmic practices, for example around the blatant biases that automated discrimination and classification perpetuates, knows ample resonances with non-algorithmic security, such as concerns on ethnic profiling. The reverse, I suggest, also holds true. It would thus be relevant analyse the role of passersby in camera surveillance as conducted by human operators, in liue of any algorithmic assessment of the subject’s movements. In their influential account on camera surveillance in public space, Norris and Armstrong (2010) describe how camera operators, informally, develop ‘working rules’ that drive their targeting. Frequently, these rules are saturated with biases. Law enforcement’s interventions ‘sanitise’ the city from those not engaging in its dominant exercise of space: consumption. Thereby, power of the watchers over the watched lies in the “promotion of habituated anticipatory conformity” (Lomell, 2004; drawing on Foucault, 1977), that is to say, the passerby adapts to the observer’s judgement. What is absent in many of these accounts, is how the routines of those passing by the camera, shape the eye and assessment of the observer.
Re-modulation
By taking a detour through technical description to understand the politics of security and threat perception in broader terms, this dissertation underscores the usefullness of engaging with security technologies at multiple granularities. For when ‘ideas about surveillance shift along with new technologies’, Galič et al. (2017) poignantly ask, ‘does that imply there are no anchor points that go beyond theories running after each technological trend?’ (Galič et al., 2017: 26) Moreover, whether developing a utopic or dystopic account of technology, making sweeping statements on paradigmatic shifts, or describing technological dumpster fires, analysis of algorithmic security politics runs the risk of essentialising the technologies it analyses (see also Aradau and Blanke, 2022). We need a politics of security technology in which critique is not simply resolved by a minor bug-fix, a patch to the assemblage, or brushed away by a benchmark.
To explore a different kind of argument, this dissertation has engaged with algorithmic technology through a hands-on technographic method of re-modulation (Annex B of this dissertation reflects on re-modulation). It is precisely the in-depth analysis of a specific technology – trendy or not – that provides insights on security politics that go well beyond the use of that specific algorithmic system. To not be at the whim of every other software update, political inquiry into security has to bolster a more thorough understanding of the devices and protocols that configure security practices. Such research is not without precedent: over the past decade ever more of such research has emerged in fields adjacent to security studies (Cox and McLean, 2013; Mackenzie, 2017; Rieder, 2020; Soon and Cox, 2020). If indeed the device’s minutae, technological nitty-gritty, matter for how environments and their threats are perceived (see also Gabrys, 2016), it means critical security studies needs to cultivate a space for research that does not shy away from these details. In a research paper, a fragment of code should be just as accepted as a quote from a practitioner or a fieldwork observation.
Fading to noise
Finally, what I hope this dissertation has demonstrated, is that a discussion on algorithmic security politics demands careful consideration of the researcher’s own positionality. All too easily, research into security and algorithmic politics places power at a distance. In these third person accounts of power, those who enact power, and those on whom power acts, remain external to the researcher. This is where multi-modal (Ragazzi, 2025), arts based (Borgdorff, 2012) research methods prove valuable. Making methodologies are not just about exploring the aesthetics of security and technology, but provide the researcher with a strategic space to iteratively configure and reconfigure themselves in relation to their object of research.
Indeed, just ‘finding out about’ algorithmic politics is not enough (Mol, 2002: 177), we need more ‘doing’, more ‘tinkering’, when it comes to a critical research into digital technologies. As explored through both the diagramming method (Section 3), as well as by making an installation (Section 5), the politics of knowledge production cannot be disentagled from the aesthetics of the artifacts by which we conduct research (see also KM Barad, 2007; Haraway, 1988b). Tinkering, thus, pertains as much to the creation of research output, as to everyday research practices. Despite decades of critique on the politics of digital technologies and online platforms, the extent to which research institutions are tied in with major platform providers is frankly staggering. Most researchers stick to their digital routines, indifferent to their own surveillance: using Google Scholar to navigate the journals, finding their way to conferences with Google Maps, hosting all their writing at Microsoft’s servers, and sending their deepest thoughts and deliberations – frequently together with someone else’s private data – to OpenAI, Anthropic and other LLM providers, who hoard this data for further model training. Despite all my work trying to avoid the platform conglomerates in my research practice33, also I plead guilty. For example, I occasionally send my recorded sport activities to Strava; an online platform where these trajectories are not just for others to see, but for the company to integrate in their models of normalcy.
As much as this dissertation is a call to others, it is a call to myself. This research reframed surveillance from a straightforward hierarchy, a one-way flow of information from the surveilled to the surveillor, to a relational account that considers the constitutive role of those who fade into the background. By gamboling about, their background noise grows louder, until all that’s perceptible is a roaring static.
Annex A Perplexity: re-modulating an anomaly detector
Documentation of an installation
The interactive installation Perplexity re-modulates an anomaly detector. Working on this installation has been essential to the arguments I develop on the algorithmic perception of movement (see Prediction of the Now p.) and suspicion (Perplexity p.).
What role do we, the public, play as collective producers of our surroundings?
Projecting pathways of light on the street, the laser light installation Perplexity interrupts the routine movements of passers-by, and invites reflection on their predictability.
Consisting of LiDAR sensors, a computer, and laser projectors, the installation reappropriates generative algorithms developed for surveillance. It tracks people passing through space, capturing their trajectories. The captured tracks are used to continuously train a model for forecasting the movements of others. These forecasts are drawn right in front of the visitor’s feet, creating an interaction with the predictive algorithm.
Perplexity explores the relations between the ordinary and the incident, between the individual and the collective, and between predictability and deviance. For, what happens if you decide to deviate from predictability?
As individuals rushing to our destinations, we may feel as if we impose our will on public space. Driven by efficiency, we chart our own route, ignore others, follow desire paths, and take shortcuts. Over time, however, these individual trajectories intertwine and connect. They collectively form patterns unique to the space. Walking the city, suggests philosopher Michel de Certeau, is inherently communicative: the movements of the pedestrian shape what is being seen and how that is understood.





Technical overview
The installation tracks movements using LiDAR sensors connected to a central computer. The computer processes passers-by’s pathways: storing them to train a predictive algorithm and using them to live forecast the passers-by’s movements. The tracked and predicted trajectories are both projected onto the ground as line animation with laser projectors. The prediction is continuously updated, creating an interaction between passers-by and the algorithm. To interject in routine movements, the installation is best placed in a transitional area, where people move with purpose – like a square or crossing.
The technical architecture is based on the principle of anomaly detection by reconstruction: an algorithm trained to predict movements serves as the baseline against which movement is assessed by comparing it with the prediction over multiple time-steps (e.g. Kanu-Asiegbu et al., 2021; Hawkins et al., 2002; Liu et al., 2018; Ye et al., 2019). The anomaly metric then is mathematically constructed as: ‖xi − x̂i‖2 (Pang et al., 2021: 38:13). This metric is also known as the L2-norm, a measure for error that describes the magnitude of difference between the two values as a distance in Euclidean space. Of the two values, xi is the observed movement (already smoothed out by the object tracker), and x̂i the prediction, both at time-step i. Movement is thus assessed at a single, distinct, time-step – a still frame from the video – that is, it is assessed as a single position in time and space. More advanced approaches, however, quantify the development of difference over time by integrating the error over the sequence in which it occurs.
Concretely, the installation relies on Trajectron++ (Salzmann et al., 2021) to forecast trajectories34. Tracking is done with ByteTracker (Zhang et al., 2022), a DeepSort derivative. Initially, the system used camera-based object detection. Therefore I trialled both Torchvision Mask R-CNN and Ultralytics YOLOv8; both trained on the COCO dataset. However, in my setup, the camera had to be mounted up high, which impeded the detection of ‘people’: the COCO dataset consists of photographs which primarily depict human figures in profile. Fine-tuning the YOLOv8 object detector on a top-down dataset only made a marginal difference. This ultimately led me to switch to LiDAR and point-cloud based clustering. Code is available at https://git.rubenvandeven.com/security_vision/trap
Listing 5: A graph of Perplexity (Note, the web version of the graph links to code).
Annex B Re-making re/search
Notes on re-modulation
In this dissertation, I have contended that we, as analysts of security, have to examine the nitty-gritty details of algorithmic systems, or we risk merely documenting whatever marketing teams dictate to be the ‘latest innovations’. Aradau and Blanke (2022) point at this tension when they write: ‘When investigating the social effects of algorithms, there is an ambiguity in the literature about whether algorithms work a little too well, thereby ushering in algocracies, or whether they do not work so well and therefore fall short of discourses of efficiency, legitimacy, accuracy, and objectivity.’ It thus is not enough to draw on promotional statements that emphasise technological change, or interpretations thereof by media outlets. Yet, reverse engineering alone will not suffice understand how algorithms ‘work’ politically (Bucher, 2018; Prophet and Pritchard, 2015: 333). While describing the distinct operations of algorithms captures their logics, it misses out on their embodied and affective effects (Cox and McLean, 2013; Ruckenstein, 2023). For that reason, this dissertation has engaged with algorithmic technology through a hands-on technographic method of re-modulation.
Re-modulation, as I discussed in the Introduction, cultivates its argumentation across the different modalities it packs together. Most prominently perhaps, the result of this approach is that the dissertation now includes an interactive installation as an integral part. While multimodal research certainly knows precedent in the transversal fields of International Political Sociology and Critical Security Studies (for an overview, see Ragazzi, 2025), it is not the most common form to present social scientific research. Over the past years I have written some words to reflect the multimodal approach to research. Yet in writing the articles that this dissertation bundles, many of these bits-and-pieces have ended up on the cutting room floor. The reflections might nevertheless prove useful for those who recognise the urgency to conduct research into the politics of technology and security through hands-on methods.
Therefore, at the close of this dissertation, I would like to spend a few pages to reflect on the methods used and how these have contributed to the arguments developed.
Allow me a preamble: I want to take care not to set such a multimodal ‘making’ method against more established, often text-based, modes of doing and disseminating research. If anything, academic writing is as non-linear as film or any other art form. A manuscript is not a still object, for integral to it are practices of writing and reading. Writing text, one seldom works their way from the text’s first word to the last, but adds words here and there, and shuffles them between sections. Moreover, the writer frequently inserts references to other works. Likewise, a reader, especially of academic writing, engages non-linearly. One starts at the title and abstract, which already provide indicators on what the text’s conclusion will be. Then, the reader skims the headers and possibly the bibliography, before settling on a point to dive in. In such writing-reading practices concepts figure as living entities that can stray well beyond the text they inhabit. Human reading and writing should thus not be considered ‘linear’, as if constituted by processing one token after the next token – similar to how large language models parse and produce text. Text-practices are not less embodied than filming or other hands-on work; they only aspire to be when advancing their theorisations in seemingly disembodied terms (e.g. Haraway, 1988b). Multimodal research explicitly does away with the dualism between knowing things for a fact and tacit knowledge (Ragazzi, 2025: 314).
The ‘re-’ is for ‘reflective’
Re-modulation is a specific kind of multimodal practice that seeks to learn about algorithmic practices by means of a modality that is integral to them. Therein, re-modulation materialises as much in the making process, as in the made object. Methodological reflection around ‘making’ or other ‘tacit’ approaches to knowledge production sometimes remains ambiguous on what the making process should do for the research35. On the one hand, the physical and emotional investment of making can guide the researcher’s thinking; the acquired “body knowledge” (Papert, 1980) is to be reflected upon and put to description. On the other hand, the emphasis of multimodal research can be on how the made object communicates an idea on registers different from text. Sitting in between these positions are pedagogies of re-making (e.g. Soon and Cox, 2020) that suggest one needs an imagined outcome to orient the making process, while it is the making itself that attenuates the researcher to particular issues. Crucially, such re-making practices seek to postpone judgements, or conclusions (Borgdorff, 2012: 71).
By denying a set end-point, re-making does not only challenge the researcher to learn about machine learning systems, and comprehend the mechanics at play, but also is a practice by which we learn from and through machine learning (Mackenzie, 2017). To do so, re-remodulation draws inspiration from critical re-modelling as it turns to a tool that is central to everyday computational practices: the program’s diagram, or schematic (Miyazaki, 2020). The ‘model’ of algorithmic operation – like a flow chart – lists the various operations a system performs: how it changes its data, which signals it sends and receives, and the conditions under which these operations occur. Re-making the algorithmic device requires the maker to trace the diagram. Re-making it “is not merely about copying a process, like a film camera would do, but furthermore is about generating new maps by slightly changing some parameters and relations entangled with such a process.” (Miyazaki, 2020: 242) In an attempt to open an algorithmic system to new interpretations, re-modulation, like re-modelling, thus actively intervenes in its model by altering the context in which it operates and the routines it invokes.
Re-modulation research is explicitly situated, and the researcher implied in the phenomena it observes (see KM Barad, 2007). The different modalities of making research challenge the researcher engaging in a technographical re-modelling to interchangeably take on various roles. Collapsed into a single being, these roles cannot be fully disentangled, and yet can be in tension, as each concerns a different predisposition – the state of mind – the researcher is in while working on the various modalities of making/research. Here, I highlight three such predispositions.
Tracing trajectories; socio-technical routines
This brings me to another role: in a re-modulation, the researcher has to place the empirics encountered into lineages and genealogies of knowledge. The cliché figure of the ethnographer – who visits a remote site, drawing out maps of a village, of homes – is useful to understand how snippets of code need to be navigated and translated. When handling the diagrams, one must remember that they – like maps – are not the territory: in service of the research, many of the complexities of the practices they pertain to are inherently abstracted away. At the same time, while reading code, one should be careful not to fall into an interpretative trap, retroactively imbuing meaning into specific lines. In its role of ethnographer, the researcher has to come to terms with the “empirical filth” (Mol, 2020) that comes into play as one ‘translates’ one complex area of research – that of programming, computer science, information theory – into another. It is all too easy for the technographer to take the moral high ground over the programmer who wrote the code that is read and invoked. However, reading the code, often many years, and many kilometres apart from when and where they have been written, the researcher should bear in mind the technical, socio-political dynamics that come into play when writing software. Not only does the code have to compile – that is, it has to be syntactically correct – it also has to solve a computational puzzle, while being constrained by everyday practicalities such as deadlines and access to resources.
While messy, broken, or outdated code tends to form a burden in
programming practices that optimise for efficiency, in a technographic
account these legacies can become a locus of inquiry (Stevenson and Helmond, 2020: 110). For
example, much of the algorithmic techniques mobilised in this
dissertation seemed to be ready for ‘off-the-shelve’ use in my
installation, integrating them often was not that easy; and
sometimes just a pain. Many algorithmic techniques have been shared as
academic papers which come with code repositories. These repositories
serve as proof of concept. Most contain some documentation: a
README.md file describes the installation of the code’s
dependencies and might showcase some commands by which to
invoke it. Reading these, a researcher (i.e. me!) might get the
impression to be up-and-running with these systems in no time at all.
This could not be further from the truth. Many of these repositories
have never been structured to be easily set up by others. Often,
software setups rely on literally thousands of small dependencies:
little (or big) software libraries each wrapping some algorithmic
technique39. As time passes – even just a year
– the code’s dependencies are deprecated: they no longer compile on
newer systems, current versions might have incompatibilities that lead
to installation or runtime conflicts. Some dependencies have been
removed from distribution networks; I often ran into this when looking
for the training datasets. Developing Perplexity deprecated
dependencies forced me to trace the code and their accompanying papers
to their logical predecessors. In many cases, it turns out, these were
simpler and yet shared their fundamental operations. While at the
surface (and in brochures) many algorithmic surveillance technologies
seem unique and novel, it is by digging into the exact techniques that a
system brings into play that one can draw out lines of coherence that
span years, or even decades of development. What is at stake in a
re-remodulation is thus not whether an implementation “works” – to
assess how well it does what it says on the tin – but to relate the
various operations that one encounters in the development process to a
larger network of techniques and practices.
Security’s participants; feeling routines
These two roles – the ethnographic-programmer and programming-ethnographer – both concern reflexive positions on making; the third role concerns that which is made. The made object is never a 1:1 replication of the device or technique that is being replicated, if we can even speak of a singular, stable original in the first place. By changing the context in which people interact with the algorithmic device, its schematic is inherently shifted. Moreover, like re-modelling, the researcher navigates these changes. By adding or altering an arrow in the system’s diagram, one changes a technicality in the algorithmic procedures, that not only bears technical consequences, but intervenes in the ways the users or subjects of the system interact with it.
Re-modulation is thus involved with the spectator, the audience, or, in the case of an interactive installation, the participant. Note that this ‘spectator’ is not external to the maker. The fuzzy distinction between the making and the made can be found in many artistic/making based approaches to research, for it is integral to artistic approaches. In some sense, the maker is configured as spectator to its own making process. Still, when it comes to interactive installations, the participant is commonly figured as an exhibition visitor that purposefully arrives at the object and engages in the ‘designed’ interaction. However, when testing Perplexity, as I set it up in the gardens of the Hof van Cartesius in Utrecht, I observed the responses of those who encountered the installation unwittingly. Those visiting their studios and offices were surprised to see the lines being drawn; they looked around, wondering what was tracking their movements; and some, for a brief moment, lingered with the installation. These reactions made me realise the common framing of a purposeful participant would be devoid of an important ingredient to the system’s operation: the routines of passers-by.
Perplexity, a re-modulation of an anomaly detector, had interjected itself into the routines those passing by. All that time, these routines had remained implicitly present in the system’s diagram. To be sure, the diagram described how trajectories of movement were recorded. Yet, the generative model that predicts the subject’s next step had been tailored to replicate the patterns these trajectories form. This patterning of trajectories – operationalising the tension between the incidental and the normal – routines become a, if not the, constitutive element of the installation’s predictive system: the passers-by, whose movements are integrated into subsequent predictions, are participants.
If we squint, we can liken the participant’s surprise to the surprise of the security operator and to the information theoretical ‘surprise’ integral to algorithmic prediction. In all these cases, the degree to which the prediction is accurate influences the surprise. Interesting to note a difference however: whereas algorithmic surprise is maximised when predictions are inaccurate, the surprise of the participant seems to be most intense when predictions are correct. This, I would speculate, cannot be seen separate from the sudden shock when realising one’s own predictability; from which it is only a small step to complacency to surveillance.
Perplexity is not an ‘object’, or a bounded installation. As a re-modulation, Perplexity spans the making and the made. By creating a space to pause, and consider how algorithms make one feel, re-modulation works by re-developing, re-tracing, and re-experiencing algorithmic devices. Doing so, Perplexity enmeshes algorithmic routines, with the routines of developers, the formal routines of security institutions, and, lastly, the routines of those observed. Folding researcher into developer into participant into observed, re-modulation explores alternative formulations of security and its politics.
Annex C Inconsistent Projections
Con-Figuring Security Vision through Diagramming
Authors: Ruben van de Ven and Ildikó Zonga Plájás.
This text is published in A Peer-Reviewed Journal About Volume 11, Issue 1, 2022: Rendering Research (pp.50-65) 10.7146/aprja.v11i1.134306.
Note on inclusion in the dissertation: In accordance with PhD regulations regarding co-authored works, the primary responsibility for this article rests with me. As first author, I was responsible for the conceptualisation of the diagramming method and the development of its software. Furthermore, I led the interviews process, the subsequent analysis, and the writing of the manuscript.
In this article, we discuss how time-based diagramming can work as an alternative method of mapping. We build on current critical literature on mapping, which has shifted the attention from maps as objects or representations towards maps as processes, thereby interrogating the politics that underpin them. We focus empirically on computer vision technologies in the field of security, through interviews with selected professionals. We showcase the possibilities offered by the method and in particular the ability to record not only the finished drawing, but also the process through which it emerges. Mobilising Lucy Suchman’s notion of “configuration”, we explore how the diagrams draw together different figures. Their loose visual language allows them to con-figure complex, sometimes even incompatible concepts and narratives in a shared visual space. Rather than forming a coherent whole, the diagrams embrace an inconsistent projection. In their unfolding over time, the diagrams forefront how such configurations are not stable structures, but rely on hesitation and contingencies.
Introduction
There is complexity if things relate but don’t add up, if events occur but not within the processes of linear time, and if phenomena share a space but cannot be mapped in terms of a single set of three-dimensional coordinates. (Mol and Law, 2002: 1)
In the exploratory phase of social scientific research, maps have often been used as valuable tools to capture, analyse, and portray an object of research. Often in the form of geographical maps, (social) network visualisations, or point clouds,40 the rendering of data points onto a two-dimensional plane can expose relations between various properties, entities, areas, clusters, or classes. However, more recent literature in science and technology studies (STS) and feminist critique of technoscience have gradually shifted attention from maps as epistemological devices to the means by which they are constituted and the politics they perform (D’Ignazio, 2020; Kitchin and Dodge, 2007). Maps are considered to have trouble addressing the fluid and messy nature of social reality while operating under a veil of neutrality (D’Ignazio, 2020; Drucker, 2011). Through their consistent mode of operation, maps perform a rhetoric "god trick of seeing everything from nowhere" (Haraway, 1988a: 581). The categories and labels of a map are no longer taken for granted, but are rather considered as a site of politics and contestation. In effect, an examination of maps is an analysis of how boundaries between entities are drawn, how differences are made, and what is included or omitted. A reflexive approach to visualizing data —explicitly or implicitly—should interrogate not only the contents of the underlying dataset, but also the way it is constituted; its structure, modes of collection (e.g. Marres and Moats, 2015; Martin-Mazé and Perret, 2021), and modes of visualisation (Dávila, 2019; Drucker, 2011). For example, by blurring lines and drawing uncertainty, a map can be more explicit about the insecurity of its categorisation (Drucker, 2011). Such a map no longer consistently projects input data onto an output surface, but instead draws attention to the practices and politics of its knowledge production.
In this text, we take these insecurities of mapping as a productive analytical site. Based on our own exploratory research in computer vision technologies in the field of security, we will outline a method that allows us to examine how our object of research emerges as a multiple, entangled in situated practices that engage with security vision.
The authors of this article are members of a research group studying the politics of computer vision technologies in the field of security. Such computer vision technologies automate the analysis of photo or video footage in order to spot weapons, violence, or other kinds of behaviour deemed undesirable, and they are increasingly being used to automate border security, contribute to smart CCTV, and moderate online conversations. In order to grasp better this field of research, we started by exploring how our object of research—“security vision”—configures notions of security and computer vision.41
“Configuration” as an analytical concept was coined by Lucy Suchman (2006) to describe how technologies can be considered assemblages of heterogeneous human and non-human elements that produce meaning as they come into relation. Suchman and other relational theorists in science and technology studies (STS) have argued that the actions of technologies cannot be ascribed to a singular actor — whether human or non-human — but instead should be considered “an effect of practices that are multiply distributed and contingently enacted” (Suchman, 2006).42 Suchman’s conceptualization resonates with Karen Barad’s notion of “intra-action” (2007) to underscore how the entities that come into relation are not given in advance, but rather emerge through the encounter with one another. What is of interest for a relational analysis is therefore not the network itself, but how such networks structure actors and entities (human or otherwise) and the complex arrangements between them (e.g. Callon, 1999). In other words, for Suchman, how humans and machines figure together or configure is not given, but rather constructed in both discourse and practice.
Fundamental to the notion of configuration is how, through the work of technologists and users, technology materializes some of the cultural imaginaries that inspire them and which, in turn, they enact into being (Suchman, 2006: 226). In our understanding, imaginaries are not the opposite of knowing or doing, but very much a part of them. These imaginaries enfold individual experience, collective professional practices, and widely circulating narratives about technology. They bring together heterogeneous elements such as one's understanding of techniques, equipment, or the juridical. Imaginaries shape and are shaped in turn by the practices of those working with technology. As such, technologies can be considered to bring together elements from across various registers into more or less stable material-semiotic arrangements. Suchman explains, “configuration in this sense is a device for studying technologies with particular attention to the imaginaries and materialities that they join together” (Suchman, 2006: 48). Configurations also draw attention to the political effects of everyday practices and how they institute bounded entities and their relations.
Taken as a site of politics, the configuration of entities is potentially an important locus of analysis. For our case, this implies that there is no single “security vision” that comprises a pre-determined set of components, but rather that such a security vision is multiple and heterogeneous. Annemarie Mol in her discussion of the ontological multiple argues that bodies, objects, and entities do not exist in and of themselves, but come into being through practices (Mol, 2002). As practices vary, so do the different enactments of the objects that are brought into being while still unified under a single nomenclature. These practices do not enact multiple perspectives on the same thing, but instead they allow a research object to emerge as more than one while being less than many. Grasping how security vision is enacted differently through different professional practices that are engaged with such technologies might help us to examine further how these technologies come to matter.
How can we then explore this “security vision” as a site that draws entities together and establishes the borders and relations between these entities?
To address this question, we mobilise the notion of con-figuration in order to propose an approach to mapping based on diagramming. Through this method, we are interested not in the finished drawings as artefacts, but rather in drawing and diagramming as time-based processes. Second, we will unpack how through con-figurations, our object of research, security vision, is rendered in spatial terms. In the third and last section, we argue that the temporal dimension within and across the various diagrams sensitises us to the uncertainties, hesitations, speculations, and inconsistencies that are instrumental in con-figuring our object of research.
Diagramming as a mapping device
Diagramming, as O’Sullivan explains, can be understood as a device that performs abstractions, suggests connections and compatibilities, and offers a perspective, a speculative future. As such, they “double as protocols for a possible practice” (O’Sullivan, 2016: 13). Diagrams historically hold an important place in computational practices (Soon and Cox, 2020). For example, a flowchart is a kind of diagram often used to describe the various steps of a programmed routine. The format of a diagram is indicative of programming as a social and communicative practice (Soon and Cox, 2020: 214). In a similar vein, in his exploration of machine learning practices, Adrian Mackenzie (2017) suggests that mathematical formulae that appear in computer science papers and software code can be seen as diagrams. Diagramming, being a spatialisation of symbols, is fundamental to computational practices. However, we propose the use of diagramming not as object of research, but as a methodological device to understand such practices.
In doing so, we take inspiration from the fields of art and design. For example, Louise Drulhe in her work Critical Atlas of Internet (2015) explores several metaphors and graphical languages that have been used to represent the Internet. The project’s loose visual language allows for the Internet to appear as a heterogeneous system of people, equipment, techniques, and material and social issues. Moreover, the various diagrams are not compatible; they are not different perspectives on the same thing. In Drulhe’s Atlas, the juxtaposition of these various renderings makes their politics visible.
The drawings bring together different entities through different relations. Seen through the analytical lens of configuration, these drawings present their object using different figures, which appear together in different con-figurations. “To figure is to assign shape, designate what is to be made noticeable and consequential, to be taken as identifying.” (Suchman, 2012: 49) Drawing a shape on a canvas is an act that draws in imaginaries in a practice of signification. Through their circulation, such figures transform as they appear in new contexts, taking on new relations and significations. The trope of the figure is suggestive of both their productive potential and the possibility of their analysis.
By taking diagrams as con-figurations, we propose a practice of mapping different from a more traditional form of consistent projections such as geographical maps. This method introduces hand-drawn mapmaking within an interview setting, allowing us to process the conversation and its image in a new way. With this method, we want to map our object of research by attending to the various ways in which “security vision” draws together different imaginaries of technology.
We conducted interviews with various professionals working in the field of computer vision and security and asked them to describe how they see computer vision operating in their specific fields. Based on an initial survey of security vision practices, in Europe we identified various roles involved in such practices. Our interviewees develop such technologies themselves, work on projects in which such software is developed, or are critical of the use of security vision, either from a legal or activist perspective.
Eventually, we conducted six in-depth interviews with professionals in three different European countries.43 Gerwin van der Lugt is a developer of software that detects so-called “high-impact crimes” in camera streams. András Lukács is a senior researcher and coordinator in the AI Lab at the Department of Mathematics of the Eötvös Loránd University in Budapest while serving as an AI advisor for the Hungarian Ministry of Technology and Innovation. Guido Delver is an engineer and coordinator of a Rotterdam-based project entitled “Burglary-Free Neighborhood” that aims at developing autonomous systems built into street lamps to reinforce public security. Attila Bátorfy is a journalist and data visualization expert who teaches journalism, media studies, and information graphics at the Media Department of Eötvös Loránd University. Peter Smith (pseudonym) is a senior security expert working for a European organisation employing border technologies. Finally, Ádám Remport is a Hungarian legal expert and activist working specifically on state actor use of biometric technologies. Being a rather small group of people, these interviewees do not serve as “illustrative representatives” (Mol and Law, 2002: 16–17) of the fields in which they work. However, as each of them has different cultural and institutional affiliations and holds a different position with respect to working with security vision technology, they cover a broad spectrum of engagement with our research object.
We began the interviews with a very basic question: “When we speak of security vision we speak of the use of computer vision in a security context. Can you explain from your perspective what these concepts mean and how they come together?” We then asked our interviewees to draw a diagram or mind map of the entities, institutions, and processes they mentioned throughout the conversation, as well as the connections between them. As the questions are asked on the spot, the con-figurations that appear can by no means be taken as exhaustive, but instead are closely tied into the conversation that brings them about.
We did not want to confine the interviewees to a particular visual register or drawing style, and nor did we want to overwhelm them with a plethora of options. Therefore, we decided on an empty drawing canvas. While we initially experimented with filming the drawing of the diagram by placing a camera above a sheet of A3 paper, we soon decided to record the drawing digitally. With a rasterised video, there would have been no (direct) way to recover individual shapes and segments in a time-based manner. Therefore, instead of using a relatively simple screen recording, we decided to develop our own software to interface the conversations.44 We wanted a vector-animation of our conversation so that at a later stage, we would be able to extract these strokes from the diagram, independently of whether they were drawn on top of one another.
In our trials, we used a standard set of four markers: black, red, blue, and green. For our digital interface, we decided to use the same set of colours. We purchased a pen display which, with its 24” diagonal, is comparable in size to an A3 paper. Mimicking the pen-and-paper set-up, we decided not to implement an undo function; instead, interviewees would have to cross out any unwanted elements of their drawing. A major difference between a sheet of paper and the digital drawing board, however, is that the latter can be dragged around, creating an infinite canvas. The diagrams that emerged through these interviews are a combination of the recorded audio with the recorded drawings, both in a time-based format.
Diagrams, O’Sullivan (2016) proposes, allow for a composite practice in which drawings from “different milieus” or frameworks can be juxtaposed as well as superimposed on one another. Such composites might help to work out possible relations and divergences among the various diagrams we collected through our interviews. Such composites could appear as collages or as time-based video edits. In this first methodological experiment, we decided to juxtapose excerpts of the diagrams using annotations as a way to have them work together.45 We therefore created an interface through which the various diagrams could be explored, taken apart, and reassembled as new wholes (see figure 13). This happens in two steps. First, we annotate the diagrams based on the conversation, the drawing, or a combination of the two. This is a rather common method for working with interviews; yet, as we work with vector animations, it allows us to extract and collect not only spoken text, but also to create excerpts of the drawings. Second, these annotations provide an entry point into the conversations; they become a way to order and see them side by side (figure 14).
With the diagramming method and tools presented here, we aim to explore the relations drawn and the entities demarcated as a way to examine how “security vision” joins them together. Diagrams as a form of mapping are exploratory devices. However, contrary to maps that serve as tools for (re)presentation, the diagrams create spatial con-figurations that do not abide by a consistent projection. In the sections that follow, we will outline how the materialization and spatialisation of the conversation that the diagrams facilitate helps us to examine the con-figurations they bring about. Subsequently, we will examine how the temporal aspect of these diagrams leaves room for uncertainties, helping us to describe how unstable boundaries solidify.
Traces of the diagrams
Before we started the interviews, we held certain expectations about what the diagrams might look like and how they would draw out various security vision configurations. The Critical Atlas of Internet, was just one of the diagramming projects that informed our expectations. Kate Crawford and Vladan Joler’s Anatomy of AI (2018) and Matteo Pasquinelli and Vladan Joler’s “spurious and baroque” Nooscope (2020b) diagram also served as visual referents when we started to develop our method. What all these diagrams have in common is that each drawing gives shape to their specific objects of research in a coherent structure. In these maps, all represented institutions, techniques, and technologies are directly or indirectly connected through the relations drawn. Therefore, we also expected that every conversation would yield a diagram that would abide by a single structure — albeit more modest in scale and more explicitly positioned than the examples mentioned. We thought that our interviewees would end up drawing circles, connected with lines and occasionally using keywords. However, as we encouraged each interviewee to use any visual expression they felt most comfortable with, the conversations yielded rather different drawings. The resulting diagrams show a rich variety, reflecting not only the divergent ideas of what it means to draw a diagram, but also what the different practitioners had in mind regarding visual representations more generally.
Figure 15: Three excerpts from the Diagrams that showcase drawings using very different visual languages. a — By András Lukács. In this excerpt, we only see bullet points with written words, b — By Ádám Remport. An illustration of a protest monitored by cameras drawn as a crowd and technological devices, c — By Attila Bátorfy. A drawing of relations between various institutions involved in security in Hungary.
This rich variety of the diagrams forced us to reconsider the conventions by which we interpreted the drawings. While the drawings often contain words, they do break with the common spatial logics of both written text and graphic design. They neither systematically flow from the top left of a canvas to the bottom right (see figure 15 (a)), nor do they present their information in a visually hierarchical way. Some of the drawings contain graphs (see also figure 17), yet they do not abide by mathematical rules. Some drawings contain arrows or lines, indicating some kind of flow or hierarchy, but these signs seldom denote clearly defined relations (figure 15 (c)). On still other occasions, relations were depicted with illustrations (figure 15 (b)). The diagrams were not clear-cut flowcharts depicting how the technology works or what it comprises.
In making sense of these diagrams, we therefore turn to the notion of con-figuration. As Suchman explains, “figuration alerts us to the need to recover the domains of practice and significance that are presupposed by and built into particular technological artefacts, as well as the ways in which artefact boundaries are naturalized as antecedent rather than ongoing consequences of specific socio-technical encounters” (Suchman, 2012: 50). The diagrams through their spatial enactments allow for security vision to emerge entangled with complex and multiple practices without naturalizing any of its terms or socio-technical arrangements. The strokes in the drawing create a material image that allows analysis of the entities it joins together, while resisting any attempt to be synthesised into one single coherent narrative.
When looking across the diagrams we collected, we can identify two characteristics of con-figurations that emerge in their visual rendering: order and multiplicity.
Spaces that order relations
First, the diagrams use space to order concepts and relations. For example, in one of the interviews, when Guido Delver discussed the “stakeholders” of the project he managed, he did not list them, but placed them instead on two axes: municipality/police ↔︎ residents and industry/suppliers ↔︎ research/universities. All these parties “gathered around” the public space in which computer vision was deployed. During another interview, András Lukács used bullet points with written concepts, but instead of placing these vertically, he placed these elements between two extremes: “security” and “computer vision”. In this drawing, security vision emerges in the centre of the image, where the two extremes overlap (figure 16, Top). On another occasion, Gerwin van der Lugt more explicitly drew a Venn diagram to locate his expertise on computer vision in a particular subset of the field (figure 17).
In the space that emerges, the placement of various concepts helps to indicate what differentiates and what unites the object of research. Through these spatialisations, we learn that the relations between the entities mentioned in the interviews cannot be reduced to either connection (as would be signalled by a line in a network visualisation) or containment (as in a Venn diagram). They are much more complex. Sometimes connections are assumed but left implicit, while at other times they are signalled only by bringing two entities into physical proximity but never spelling the connections out. Connections are made explicit only when they figure in a specific story line.
Multiple configurations
The second way in which space matters in the diagrams is to allow for multiplicity within the drawings. Most drawings, while forming a whole within the context of the conversation, can also be seen as being composed of many distinct drawings that are the results of loosely connected topics discussed by the interviewees. These distinct drawings appear side by side, sometimes even curving around one another, ever shifting in scale. Scribbling asides in the corner, the interviewees often tried to squeeze as much as possible within the boundaries of the 24”-inch canvas. The equations in figure 17, for example, were so squeezed in that they had to be explicitly demarcated from the rest of the drawing by a line. Even though the interface allows for infinite dragging and is theoretically unbound, the thick black borders of the pen display did in fact matter in shaping the drawing, as the drawings try to take up the space that is left available to them. The absence of a uniform projection liberates these multiple drawings-within-a-drawing from a mutual visual hierarchy. While this might make the reading of the diagrams difficult, it allows the diagrams to bring together concepts and visual language from across various incompatible registers. They appear not as coherent narratives, but as collections of figures, thoughts and associations, summarising and synthesising larger ideas that hang together by virtue of their mutual appearance in the diagram.
By allowing for both order and multiplicity, the diagrams con-figure incompatible concepts and narratives. Moreover, the spatialisation of the conversation cannot be seen as distinct from the diagram’s temporal dimension. During the interviews, the drawings often became a visual referent that facilitated further elaboration and explanation. In these moments, the strokes on the canvas provided landmarks for the conversation. This becomes apparent within the conversation as the interviewees turn to the drawing to point out what they are speaking about or to pick up the conversation from a particular point. The drawings also served as visual references in the phase of their analysis. After we had conducted the interviews, we printed out the drawings on A3 paper and hung them on our office walls. While we thus temporarily flattened the diagrams, removing their temporal dimension, it was by looking at the printouts that we could retrace the conversation and recall the topic being discussed with each specific shape. The diagrams again spatialised the conversations, this time those we were having among ourselves in the office.
In the next section, we will elaborate how the diagrams, through both their temporal unfolding as well as their mutually coming together, foreground the ways in which security vision con-figures uncertainties and hesitation.
Contingent diagrams
By recording the interview as unfolding in time, the diagrams gain a temporal dimension. This allows us to see what happens before or after a stroke. When playing back the recordings, it quickly becomes clear that the drawing as a device shapes the diagrams. It pushes itself forward in the sudden line breaks when the pen is not properly touching the surface of the display; in the confusion of how to “move around” the canvas; in our requests to use different colours; in the moments when a slight hiccup in the Internet connection causes the interface to require a “refresh”. Such moments punctuate the conversations. However, when we look at the temporal unfolding of the diagrams, another kind of interruption also becomes visible.
During the conversations, our interlocutors frequently voiced doubts and uncertainties before putting the pen to the canvas. The uncertainties expressed were, for instance, about the terminology, the parties involved in a project, or relationships that “might be possible”, but whose actual status is unclear to the interviewee. Sometimes such doubts lead to crossed-out text, different line styling, or clear question marks. For instance, Ádám Remport wanted to depict a database of facial image data (figure 18). He began by drawing a collection of facial photos, at one point realising that “this is not what a facial database looks like.” He crosses out the drawing and draws another representation in which the face is “coded” instead of pictorial. The drawing and subsequent redefinition draws attention to the dominance of particular images and imaginaries of technology over others. An expert intuitively defaults to such imageries, but then feels the need to explicitly distance themselves from them. Another such moment can be found in the diagram of Gerwin van der Lugt. When he discusses the equations for true and false positive rates (TPR/FPR), he corrects his definition in the drawing while stressing the importance of being precise about these terms (figure 19). In these cases, the very act of drawing triggers hesitation and redefinition.
However, after the pen touches the canvas, the only remaining trace of hesitation is often the brief increase in the interval between strokes. The drawing solidifies the entities mentioned, even if the doubt is verbally expressed. The canvas as a medium forces the speaker-drawer to make a decision as to how to represent the uncertainty. Deliberate or not, a moment of prioritization takes place. While the interviewees give air to many of their considerations, only seldom do they choose to “give ink” to them too. It is for precisely this reason that we do not disconnect the visual from the auditory or the drawing from speech. While each track can be informative on its own, it is in the resonances and dissonances between the two (the drawing and the sound) that the diagrams allow the fuzzy nature of that which is figured to step forward. As Drucker (2011) argues, by allowing for such complexities, we can work with notions that are co-dependent and contingent without reducing or purifying them (see also Law and Mol, 2020). The act of doubting, ever present in the diagrams, blurs the boundaries of the concepts mobilised, alerting us to complexities that otherwise would be “cleaned” out.
This becomes even more apparent in the process of annotation. Through the need to provide an in point and an out point for each annotation, the conversation pushes back. When does one cut the continuous flow of a conversation in which what is being said is always in relation to what comes before and after it? Nevertheless, annotating the diagrams helps make sense of what and how security vision is con-figured by looking not only at a single conversation, but also across the various diagrams. The annotations allow us to cut up the diagrams and reassemble them into new collections. When juxtaposing these dismembered parts, we see variations appearing across the interviews.
In juxtaposing these excerpts, the diagrams remind us that they do not present absolute truths. Instead, they provide a glimpse into how our interlocutors understand and work with security vision. As such, any description counts. While one of them (a software developer) lists particular local “security integrators” as key partners in the deployment of their technology, another interviewee (an activist) considers the technology provided to governmental organisations by big tech companies such as Google and Facebook to be a threat. As configurations join together imaginaries and materialities, we need to take uncertainties and speculation seriously. Speculations abound as to which companies are involved, which technologies are used, or which futures this entails. Collaborations and conjectures, specificities and grand narratives appear side by side. Different entities configure security vision through different relations.
It is by caring for instead of rejecting these contradictions and convergences that we can get a sense of the politics of security vision that materialises between the various fields and professional practices and between the diagrams.
Conclusion
Although a single simplification reduces complexity, at the places where different simplifications meet, complexity is created, emerging where various modes of ordering (styles, logics) come together and add up comfortably or in tension, or both.
In this article, we discuss our use of diagramming as an alternative means to map the field of security vision. In an effort to account for the situated nature of the mapping exercise, we did not define security vision beforehand, but instead delegated this task to various professionals working with computer vision technologies in the security field. The resulting diagrams thereby situate our object of research in various practices such as those of software developers, engineers, program coordinators, activists, etc. The diagrams—specifically, the discrepancies and incongruities within and between them—demonstrate that we can effectively explore the con-figuration of entities and the relations among them without necessarily flattening or cleaning them, such as would happen in a straightforward visual projection.
Although we should be careful not to fetishize the affective quality of a hand-drawn diagram as opposed to that of a computer-generated map, their “sketchy” nature suggests their status as a conceptual aid. The diagramming therefore becomes “a strategy of experimentation that scrambles narrative, figuration—the givens—and allows something else, at last, to step forward. This is the production of the unknown from within the known, the unseen from within the seen” (O’Sullivan, 2016: 17). Like maps, diagrams can serve as exploratory devices. Instead of adhering to a consistent projection of in point to out point, they rely on “speculative geometries” and “self-organizing forms” (Soon and Cox, 2020: 221). Similarly, the diagrams we collected do not curate a clearly structured set of devices, institutions, or people. Rather, it is by collecting and combining a variety of diagrams about security vision that our object of research emerges as an ontological multiple. Inspired by diagramming projects such as Anatomy of AI or Nooscope, which address the politics of artificial intelligence through single visual objects, we experimented with a disjointed kind of diagram. While seemingly similar in nature, the goal of time-based diagramming is different from these meticulously designed structures. Rather than a device for presentation, the method rather helps us to analyse the structuring networks of associations.
Their loose visual language allows the diagrams to con-figure
complex, sometimes even incompatible concepts and narratives in a shared
visual space. In their unfolding over time, the diagrams forefront how
such con-figurations are not stable structures, but rely on hesitation
and contingencies. Their use of space on the canvas is no longer
consistent. As the drawing unfolds, one can see the space grow and
shrink, transforming from a two-dimensional plane into a
three-dimensional space, or even being suspended altogether. This fluid
topology opens up intriguing avenues for exploring computer vision
technologies in the field of security and locating their politics in
unexpected entities and relations (to be continued in (van
de Ven et al., 2026 bundled as sec. 3)).
Bibliography
Curriculum Vitae
Upon obtaining his atheneum diploma in 2011, Ruben van de Ven began his academic career in Mechanical Engineering at Twente University. Receiving his propedeutic degree he transitioned to Audio Visual Media at the University of Arts Utrecht (HKU). He earned his BA in 2013 directing the short film Days of Water, and an MA in Creative Design for Digital Cultures from The Open University (UK) that same year. During this period, he worked in software engineering, developing database and recommender systems for a Dutch vacancy platform. Eager to explore crossovers between film and code, between visual and digital, he joined the Media Design MA programme at the Piet Zwart Institute, Rotterdam (NL). He graduated in 2016 with artistic research on emotion recognition software.
His hands-on, visual practice has been supported by various grants and residencies. Key projects include Emotion Hero at Arquivo237 (Lisbon); the Plotting Data project with Cristina Cochior (2018); the Accept & Work series (2019–2021) with Merijn van Moll.
In 2021, Ruben joined the Security Vision research group at Leiden University’s Institute of Political Science. Supervised by Francesco Ragazzi, and with further guidance of Winnie Soon and Daniel Thomas, Ruben studied the politics of algorithmic surveillance and movement tracking. His multi-modal dissertation integrates academic writing with an interactive installation and a toolkit for visual interviews. Furthermore, he contributed to an EU-Greens/EFA report on biometric camera surveillance in Europe, and co-organised the Netherlands DoingIPS seminar series 2023, and the workshop Motion Technologies at EISA-PEC 2024.
Ruben presented his work in spaces as diverse as: V2_ in Rotterdam, ZKM in Karlsruhe, MuseumsQuartier in Vienna, State Festival in Berlin, Codes & Modes in New York City, the ArtScience Musem in Singapore, Dutch Design Week in Eindhoven, Video Vortex in India. He gave workshops at institutions such as Aarhus University (DK) and the Royal Academy for Fine Arts (NL). His research is published in peer-reviewed journals, including Environment and Planning D, International Political Sociology and A Peer-Reviewed Journal About.
Full portfolio: https://rubenvandeven.com.
Notes
In the Netherlands, the first municipal CCTV camera system was put in operation January 1999, in Ede.↩︎
I am certainly not the first to trace contemporary technologies to that point in time. For many scholars, the development of information theory and cybernetics serves as a historical reference point to highlight the disembodied ideologies of knowledge that undergird present-day data practices Hayles (1999).↩︎
More on how the tool was developped in (van de Ven and Plájás, 2022), which is bundled as Annex C.↩︎
Observed static should not be confused with stillness. Imagine listening to a detuned FM radio; set not to receive a station, but amplifying frequencies lying in-between. All you hear is static. This static, “white” noise is audible (hissing, crackling) due to the movements of various sources of electric charge. Static is a disorganised ‘nothing’, that is only amplified by the radio’s circuitry in the absence of a modulated signal. Static moves erraticly. While these movements are not predictable, they adhere to probabilities. Static thus does not contain information, for it does not yield surprise. Rather, static is the result of an ethico-onto-epistemological process of backgrounding and foregrounding movement.↩︎
For example, ViNotion speaks on their website about “smart camera sensors” (ViNotion, 2023); see also van de Ven et al. (2026) (bundled as sec. 3).↩︎
What we thus see at play is similar to the logic of the operational image (Farocki, 2004): unpacking a
camera + computer = sensorreveals a whole cascade of signals, appearing as ephemeral numbers, meant to be discarded before they are looked at by an end user.↩︎Nor, on the other hand, does it mean that asylum decisions are entirely at the discretion of the caseworker↩︎
As noted in the conclusion in this collective discussion piece, in particular in relation to human mobility and migration, the binary of perceptibility/imperceptibility in relation to political struggle should be problematized further.↩︎
The published paper contains the following note on funding: Research by Klaudia Klonowska was conducted within the framework of the DILEMA Project, which received funding from the Dutch Research Council (NWO) Platform for Responsible Innovation (NWO-MVI, Grant No. MVI.19.017). Sofie van der Maarel’s research was funded through the NWO-NWA project “Understanding and preventing moral injury among military and police personnel” (Grant No. NWA.1160.18.019). The work of Ruben van de Ven was supported by funding from the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation programme (SECURITY VISION, Grant Agreement No. 866535). Jasper van der Kist’s research was supported by the Fonds Wetenschappelijk Onderzoek (FWO) project “Platform Wars” and Riksbankens Jubileumsfond (Award No. P20-0618).↩︎
The interface software and code is available at https://git.rubenvandeven.com/security_vision/svganim and https://gitlab.com/security-vision/chronodiagram↩︎
The study strictly followed the guidelines of the European Research Council in terms of informed consent, privacy and data retention, as specified in the ethics deliverables of project 866535, SECURITY VISION.↩︎
See https://www.securityvision.io/diagrams/web/companion.html↩︎
We would like to thank the reviewer who pointed out the possible links between these two logics.↩︎
Ádám Remport is a Hungarian legal expert and activist working on state actors’ use of biometric technologies.↩︎
Gerwin van der Lugt is a developer of software that detects ‘high-impact crimes’ in camera streams.↩︎
Sergei Miliaev is principal researcher and facial recognition team lead at VisionLabs in Amsterdam.↩︎
András Lukács is professor of artificial intelligence and data science at the Department of Computer Science, Eötvös Loránd University, Hungary.↩︎
John Riemen is head of the Center for Biometricts for the Dutch police.↩︎
Guido Delver is an engineer and coordinator of the Rotterdam-based Burglary-Free Neighbourhood pilot-project, that builds autonomous systems into street lamps to reinforce public security.↩︎
Dirk Herzbach is chief of police at the Police Headquarters Mannheim, Germany.↩︎
Jeroen van Rest is a safety expert and senior consultant in risk-based security at TNO, the Netherlands.↩︎
In this text, I use perception to denote what is technically known as ‘transduction’: the active process by which sensor input is converted into a signal (van der Maarel et al., 2026).↩︎
The distinction between ‘real-time’ observational practices and pre-emptive security is not absolute. In a naive understanding, real-time observational practices monitor the present, while the pre-emptive governance relies on statistical techniques to make unlikely futures actionable. Hence, real-time surveillance has been contrasted to ‘sorting’ to distinguish the immediacy of watching eyes from anticipatory security practices (Fuller, 2007: 149; Lyon, 2003). In contemporary surveillance practices however, the distinction is not clear-cut. First, real-time practices are ever near real-time. Delays incurred by data transfer and processing—which occur even at the microscopic level as light travels from the subject to the observer—means that these practices monitor movements that have passed. The need for acceleration of processing is fuelled by the desire to decrease the delay, the latency, between an event and the moment of observation. This plays, for example, at FRONTEX: an ever more complex assemblage of communication devices is deployed to improve ‘situational awareness’ of the authorities (Dijstelbloem et al., 2017). By channelling observation data from border zones to centralised control rooms at increasing speed, it is believed that governing bodies gain a more complete ‘situational picture’, by which they then can prioritise their interventions. In the production of this situational picture, near-real-time surveillance is complemented with a logic of risk, for such an image not only presents present movements, but, in the process of prioritisation, relies on calculations that quantify potential futures, anticipating imminent harm (Walters, 2017: 807). In both accounts, anticipatory logic is used as a method of forecasting, of grasping unknown futures by means of probabilities, to direct action in the now.↩︎
If we merely consider the temporal effects of technology in terms of its velocity, of its more rapid data processing, we risk buying into a linear timescale that is entangled with progressivist and productionist understandings of innovation and technology (Puig de la Bellacasa, 2015; see also Kitchin, 2014; Hopman, 2024). Aradau and Blanke (2022) point at this tension when they write: ‘When investigating the social effects of algorithms, there is an ambiguity in the literature about whether algorithms work a little too well, thereby ushering in algocracies, or whether they do not work so well and therefore fall short of discourses of efficiency, legitimacy, accuracy, and objectivity.’↩︎
The categories available in many software packages can be traced to some canonical training datasets used for machine learning such as COCO or Pascal VOC (Cochior and van de Ven, 2020).↩︎
The code of the original DeepSORT implementation is available at https://github.com/nwojke/deep_sort↩︎
On the notion and methodology of technography, see Bucher (2018). The code for this paper’s visualisations is available as supplementary material and at https://git.rubenvandeven.com/security_vision/mot-paper-examples.↩︎
The text is automatically colour coded by the code editors that programmers tend to use. The colour has no intrinsic meaning, but highlights the syntactical category of a term. This makes it easier for a programmer to recognise the operations that the code describes.↩︎
In some programming languages, the for-loop describes a finite operation to loop over for example a given list. In the code here, written in Python,
videois a ‘generator’ that can provide frames for an unspecified — sheer infinite — number of iterations, for example, coming from a live network surveillance camera.↩︎As an example of why one cannot always measure exactly what one is interested in, State observers | understanding kalman filters, part 2 (2017) provides two examples: first the emotional state of an individual – which is non-numeric anyhow – and the temperature of rocket fuel, which is so high any sensor would melt. Thus instead of measuring the internal temperature, one can measure the external temperature, and use that to estimate the internal temperature.↩︎
In the original DeepSORT implementation, all features associated with a track are stored and used to associate incoming detections. In more recent variations on this technique (i.e. Stanojevic and Todorovic, 2024; Aharon et al., 2022; Wang et al., 2020), these identifying features of the object are integrated into a single descriptor which, again, mitigates noise in individual assessments:
↩︎Such a re-consideration of camera surveillence beyond the use of algorithmic processing is all the more relevant in a Dutch condext. Last year, the national command of the Dutch police – who in the Netherlands oversees municipal CCTV – banished any trials with algorithmic analysis of behaviour in municipal camera surveillance. The decision came top-down, and surprised many involved in the trials. While the exact reasoning behind the decision has remained classified, some involved informally say it had been out of fear for ethnic biases these systems might, or might not, perpetuate.↩︎
The members of Security Vision, the research group of which I have been part for the past five years, have experimented with alternative infrastructures. We self-hosted our documents with NextCloud, did our collective writing with OnlyOffice, and I hosted the code I wrote at my own server. But surely, this was not a seamless experience. As I write these words in my text editor, following Markdown formatting, I already dread the moment I need to manually copy-and-paste suggested revisions, instead of just pressing ‘accept’. Imagining alternative infrastructures requires time and deliberation.↩︎
Some changes have been made to the codebase of Trajectron++ to facilitate its use as a module in Perplexity’s codebase, and to speed up online inference. See https://git.rubenvandeven.com/security_vision/Trajectron-plus-plus↩︎
Telling herein are discussions about ‘critical making’. This term coined by Matt Ratto (Ratto and Hockema, 2009) and developed together with Garnet Hertz (Hertz, 2012). While both acknowledge the embodied and affective dimensions of multimodal research, their emphasis on when making plays a role differs.↩︎
See https://git.rubenvandeven.com/security_vision/ .↩︎
It is quite worthwhile to read how Redmon, himself funded by the U.S.A. Office of Naval Research and Google, puts it in the paper with which he published YOLOv3: ‘But maybe a better question is: “What are we going to do with these detectors now that we have them?” A lot of the people doing this research are at Google and Facebook. I guess at least we know the technology is in good hands and definitely won’t be used to harvest your personal information and sell it to…. wait, you’re saying that’s exactly what it will be used for?? Oh. Well the other people heavily funding vision research are the military and they’ve never done anything horrible like killing lots of people with new technology oh wait….. I have a lot of hope that most of the people using computer vision are just doing happy, good stuff with it, like counting the number of zebras in a national park, or tracking their cat as it wanders around their house. But computer vision is already being put to questionable use and as researchers we have a responsibility to at least consider the harm our work might be doing and think of ways to mitigate it. We owe the world that much.’ (Redmon and Farhadi, n.d.: 4)↩︎
Running
uv pip freezeon the trajectory predictor (‘trap’) code reposity gives 217 dependencies, andcargo treeon the laser projector part of the installation (laserspace) yields 1702 dependencies.↩︎For instance, when using Multiple Correspondence Analysis (Le Roux and Rouanet).↩︎
Note that this paper is a methodological exploration. An analysis of “security vision” through the lens of diagramming is featured in another article (van de Ven et al., 2026 bundled as sec. 3). For an discussion on the exploratory phase, see Plájás et al. (2020) .↩︎
In a famous example, Bruno Latour describes how it is neither a gun nor a human individual that shoots (and, in effect, potentially kills), but instead the act of shooting is mutually constituted by both human and non-human actants: “You are different with a gun in your hand; the gun is different with you holding it” (Latour, 1999: 179).↩︎
The Netherlands, Hungary, and Poland.↩︎
The code for the interface is available at https://git.rubenvandeven.com/security_vision/svganim↩︎
Other ways of working with the diagrams could prove interesting as well. For instance, we have considered overlaying handwritten annotations on top of the diagrams. Another possibility would have been borrowing techniques from qualitative interviewing: we can visit the interviewees several times, each time refining the drawings, or discussing the diagrams of other interviewees to elicit additional reflections on, or reconfigurations of, their initial input.↩︎