Through the Praxis Lens
We need to have ethnographers in AI labs
Contents
Abstract
The fast-paced design, development, and deployment (DDD) of AI systems across a range of environments speak to a need to better understand the context in which these systems are created and used. This white paper proposes a conceptual framework of ethnographic studies for capturing the diverse sociocultural contexts and human practices around the entire lifecycle of AI systems. It outlines a systematic approach to embedding ethnographers throughout the AI DDD process. Drawing from existing literature on ethnographic and Participatory Design (PD) studies, this paper offers a breakdown of the AI mega-process (research, design, development; testing and evaluation; deployment and iteration) and suggests positioning ethnographers next to participatory designers for a deeper understanding of these mega-processes. As such, it helps make AI developers’ and marginalised groups’ practices visible, thereby accounting for a full picture of AI DDD. While emphasising ethnographic and PD studies in AI DDD, this paper recommends further research to produce case studies that incorporate embedding ethnographers throughout the AI lifecycle.
1. Introduction
“Innovation is an imagination of what could be based in a knowledge of what is.”1
The design, development, and deployment (DDD) of artificial intelligence (AI) systems2 are taking place at an unprecedented scale globally, with significant AI research and development (R&D) activities growing exponentially from 2011 to 2023 (Maslej et al. 2024). However, this fast-paced development and adoption of AI systems is not yet a universal phenomenon; certain parts of the world like the United States, China, European Union countries, and the United Kingdom dominate in AI R&D, while the rest of the world is left to play catch up (Maslej et al. 2024). On the one hand, there is the need to level the playing field so everyone can benefit from AI systems. On the other, the larger issue is that AI systems created by and for a handful of countries are increasingly exported to be implemented in other societies of drastically different economic, political, and sociocultural environments. This clash of environmental contexts — the difference between the one where AI systems are developed and the other where AI systems are deployed — is not obvious, but social science researchers, anthropologists, and even some AI developers have begun to take note.
For example, the large language models (LLMs) that took the world by storm in 2022 were reported to exhibit cultural biases (Liu et al. 2024; Cao et al. 2023; Masoud et al. 2024). Researchers further found that a diversity of LLMs reflected ideologies associated with the countries and cultural environments where they were created, showing that developers’ design choices from compiling training corpus and finetuning to system prompting affect the behaviours of AI systems (Buyl et al. 2024). For example, an image recognition system designed by Google AI in the United States to assist Thai health workers’ diagnosis procedure could not handle the quality differences between training data and inference data, the technical constraints imposed by Thai clinics’ equipment, and the discrepancies in imagined workflow of health workers (Beede et al. 2020). In the end, the system disrupted the workflow more than it helped it. Yet another, when Indian users interacted with the algorithmic-driven digital dating platform created in Silicon Valley, they reported experiencing clashes of values, as the underlying logic of the platform design differed from their daily experiences embedded in Indian sociocultural backgrounds (Kuthiala and McBride 2025).
These cases reveal a few issues. First, the sociocultural environments in which AI developers are embedded and their particular practices, design choices, and personal worldviews may all affect the design and development of AI systems, but most of these factors are hidden behind the wall of proprietary intellectual properties, away from the public’s view. Second, AI systems created solely based on AI developers’ own context (e.g. following the I-methodology) may not capture the full complexity of the diverse environments where they are finally deployed. Third, end users may have unforeseen practices, practical knowledge, and even expertise about the environment where AI systems are deployed. There needs to be a different approach to AI DDD where context-specific knowledge and expertise are made aware, and hidden human practices are made visible. With an increase in AI DDD, they must be grounded in the real world and based on how heterogeneous humans embedded in diverse environments do things.
Participatory design (PD) has been mentioned as a research methodology that can address the third issue — focusing on bringing forth the needs, desires, and perspectives of users, particularly those historically marginalized in AI system design (Gill 1989; Zytko et al. 2022; Birhane et al. 2022). While this paper acknowledges the multiple merits of PD, it argues for adding a distinct but linked theoretical orientation–namely, analytical ethnography–to understanding human practices and the context in which they take place (Anderson 1994). Analytical ethnography (hereinafter, ethnography) observes contextual human practices not just limited to end users of AI systems, but also all actors such as developers, designers, and auditors (or system evaluators), involved in the entire lifecycle of AI systems. With ethnography, all three issues of the clash of contexts can be comprehensively captured, and the humans involved in embedding such contexts into AI systems can be made accountable.
Recognising the value of PD, this white paper proposes a complementary and conceptual framework based on ethnographic approaches to account for the clash of contexts through the praxis lens during the AI DDD processes. This paper outlines the role of an ethnographer and the various research venues they can pursue across the stages of the AI DDD process. Section 2 juxtaposes the offerings of ethnographic and PD approaches and discusses ethnography’s specific contribution to AI DDD. Section 3 introduces and outlines the conceptual framework for embedding ethnographers at each stage of the AI mega-process (research, design, development; testing and evaluation; deployment and iteration). The paper concludes with recommendations on the conceptual framework’s implementation and pathways for further research.
2. Incorporating ethnography in PD for AI DDD
The call for embedding ethnographic approaches with design methods is not new. In the early 1980s, computer technologies were exported from research labs and engineering environments into mainstream offices, manufacturing, educational settings, and eventually people’s private homes (Blomberg, Burrell, and Guest 2002). Designers and developers could no longer rely on their assumptions and imagination of ‘the user’ while designing computer systems but had to expand their definitions to move beyond technical experts to include non-technical laypeople who used these technologies (Blomberg, Burrell, and Guest 2002; Blomberg and Karasti 2012). This move signalled a shift towards understanding the diverse realities of people who used these tools in different ways (Blomberg et al. 1993).
Applied to AI systems, ethnographies and PD both, focus on paying attention to human interactions with a technological artefact and the context in which it happens; however, their theoretical orientations, approaches, and focus differ. First, ethnography and PD look at human practices and knowledge with different intentions. PD involves and morphs the users’ voice, knowledge, and perspectives in the design and development process of a technical system (Delgado et al. 2023). On the other hand, ethnographers often do not start with the premise of designs; instead, they focus on comprehending the “use-contexts” of designs, documenting current and real-life settings of human practices (Blomberg, Burrell, and Guest 2002). In other words, PD has a future orientation to always look at user practices and preferences through the lens of what can be designed, and ethnographies focus on understanding the current “enacted local practices” (Crabtree 1998)).
Second, PD’s focus on users’ voices and knowledge grounds technical design in what people say they do. However, ethnography acknowledges that what people say and what people do can be very different. Ethnographies capture this nuance by providing detailed accounts of people’s real-life practices and behaviours through the interpretive eyes of participating observers (Blomberg, Burrell, and Guest 2002). In other words, an ethnographer’s attention in a study is not confined to what is said but rather focuses on what is observed and noticed, opening up to diverse possibilities of practices that may emerge in the complex realities. This allows ethnographers to contribute to designers by “opening up the overall problem-solution frame of reference,” introducing analytical sensibilities, and enabling both designers and participants to challenge their taken-for-granted assumptions about what is and what could ever be (Anderson 1994).
Another distinct difference between ethnography and PD is that ethnography is not confined to observing just the end users or affected communities of AI systems, which is the core offering of PD. An ethnographer’s “motivated looking”(Anderson 1994) can capture the broader lifecycle of AI systems, putting the very design and development processes under inspection as well. This flips the script that developers and designers usually play in AI DDD: they are not merely experts with technical knowledge and neutrally creating AI systems in a sterile lab; instead, they themselves are embedded in particular political and sociocultural environments, hold personal aspirations and worldviews, and engage in a complicated web of “social and material practices” as they assemble different parts of technical artefacts to create an AI system (Seaver 2013; 2017; Wu 2024).
Ethnography is very suited to unpack the entire process by which AI systems come to be and the complex realities in which it is embedded throughout their lifecycle. This method of inquiry debunks the idea that humans who interact with AI systems are a homogenous group with similar practices, values, behaviours, attitudes, and interactions with AI systems (Oh et al. 2020) By focusing on the detailed practices, habits, values, and sociocultural norms people follow, ethnography allows a better understanding of how AI systems are created and used, and what social, political and ethical issues they raise.
3. Embedding Ethnographers in AI DDD
To address the gap in understanding the diversity of human practices in AI DDD and to contribute to existing discussions on using ethnography for innovation design, this paper proposes embedding ethnographers throughout the AI DDD lifecycle to highlight the existing practices, local knowledge, and the different values of stakeholders that inform existing AI DDD process. The ethnographers would carry out their research not as temporary external observers, but as active participants during the model lifecycle, existing for prolonged periods within AI labs3 or other relevant field sites (Table 1):
| Field site | Research examples | Relevance |
|---|---|---|
| AI labs, RD&D sites | Focus on team dynamics, decision-making, and practices within labs, uncovering biases and aligning workflows with ethical goals. | Increasing the transparency of AI design and development practices and cultivating a culture of reflection on societal impacts within AI development spaces. Permitting outsider insight into the chain of command to understand how and why AI systems are made. |
| Training and evaluation sites | Study the practices and assumptions of third-party auditors and their implications for fine-tuning the AI systems. | Fostering reflexivity on the neglected ethical or best practice requirements. Understanding how the political and cultural views affect the implicit risk hierarchies that AI systems are primarily tested against. |
| End-user contexts | Explore human interaction evaluations (Ibrahim et al. 2024) or examine AI tools in real-world use analyzing usability, cultural fit, and unintended impacts to better align systems with user needs. | Gaining insights and diverse feedback for more inclusive model design, as opposed to relying on assumptions and hypotheticals. Bridging design and deployment, therefore, provides a more holistic view of the AI supply chain. |
(Table 1. Example field site typology for studying AI DDD.)
Primarily, ethnographers can help understand the particular activities many actors (e.g. developer, designer, auditor, user)4 engage in at different stages of AI systems’ lifecycle and the dynamic interaction these actors have with each other, the technical artefacts, and society at large. Figure 1 features a simplistic conceptualisation of the AI lifecycle and its most common actors:
Ethnographer’s lens
(Figure 1. Processes and actors in AI systems’ lifecycle.)
It is critical to note that Figure 1 starts with the simplest assignment of actors to each stage of the AI lifecycle. In reality, and as will be shown in later figures, the boundaries between the roles that actors play can be blurred (e.g. users and designers may co-design during the R&D stage; users, auditors, and designers collaborate in producing future iterations of the system through feedback loops, testing, and evaluation, etc.). The ethnographer’s role, centred in the middle, is precisely to observe as each type of actor engages with each other across different stages, potentially crossing the boundaries of their assigned roles, stages and field sites.
Ethnography’s emphasis on reflexivity and its attentiveness to social relations and interactions can further uncover the social nature of AI development. As a result, this approach helps increase the transparency of the processes involved in AI lifecycles. Ethnographic tools, such as interviews or field notes, also constitute records documenting labs’ practices and the technical artefacts already significantly changing human realities. As such, these tools may help popularize the knowledge about AI development and keep labs accountable.
Ultimately, the insights produced by ethnographers can be used to productively inform the research, design, development, and testing of current and future AI systems. On the one hand, this approach heavily leverages the ethos of Human-Centered Design (HCD) and PD in that the attentive discovery about humans and society should feed important findings back into designs and development, effectively creating a tight feedback loop where meaningful actions can be taken and implemented in future iterations of AI systems. On the other hand, this framework emphasises the need to substantiate the AI lab’s development and testing processes with the ground realities that different user groups experience. As a result, it strengthens the inclusivity in AI development by bringing to light the participation of multiple stakeholders involved in the AI lifecycle, especially the marginalised actors.
This paper broadly groups the five processes of an AI system’s lifecycle (Figure 2) into three “mega-processes”:
- Research, design, and development (RD&D)
- Testing and evaluation (T&E)
- Deployment (and iterations)
In each mega-process, an ethnographer performs different analyses, visits different field sites, and creates different feedback loops to inform the next iteration of an AI system. In the long term, this knowledge production could help reshape AI DDD practices into being more societally adaptive and beneficial.
3.1. Research, design, and development (RD&D)
About this mega-process
During the R&D process conducted in labs, researchers undertake preliminary analyses of the problems their AI system is meant to solve (i.e. the objectives given to the models). This process requires them to imagine potential use cases, prospective user groups, and how their models fit into the picture. They may use user research via surveys, user group studies, or more directly base it on their personal experiences. They may also build prototypes or proof-of-concepts at this stage. In short, R&D is when researchers translate real-life or non-technical problems into technical ones that can be addressed through a technical solution.
The development process transforms this idea into a scalable product. It includes refining the prototypes and proof-of-concepts, eventually developing a feasible technical solution. It is important to highlight that some cases of AI development may not follow the same straightforward model as DDD. That is because some systems may not be generated from scratch but rather fine-tuned from other systems—building custom GPTs for specific tasks on OpenAI’s ChatGPT is an example.
Ethnographer’s lens
(Figure 2. Mega-process of RD&D.5)
Observables
During this mega-process of RD&D (Figure 2), two actions are worth observing and contesting: the translation of problems, and the corresponding solution development. Researchers and developers often have to make assumptions about the environment and the users their model will face, creating an abstraction of reality when attempting to translate a non-technical problem into a technical one. Translation can happen as they perform data annotations, make assumptions about the correlations and causality between features and targets, or fine-tune hyperparameters according to perceived suitability in a use case. Translation takes place in their daily practices, handling datasets and scripts, client meetings, brainstorming sessions, or even marketing workshops. Translation can be done by the researchers or developers, colleagues they chat with, experts they consult, executives with broader visions, or users surveyed.
In solution development, researchers and developers need to decide what works and what doesn’t. They may need to break down the defined problems into parts they can solve, focus on features that correspond to market demands, create a solution that aligns their company’s path-dependency, or prioritize cost-efficient designs. Ultimately, RD&D is composed of many social interactions and choices that are shaped by the assumptions and ideologies of the involved actors (Buyl et al. 2024). For example, creators of AI systems inherit the “Inventor’s Bias” (Cratsley and Fast 2024), which is the propensity for inventors to be over-optimistic about the positive features and uses of the products they create, such as algorithms and software. Ethnography as a method could highlight these seemingly invisible social forces, allowing developers to question their taken-for-granted assumptions embedded in conventional problem-solution design frameworks. Additionally, this would allow other stakeholders to critically challenge AI labs’ processes and encourage more participation from diverse actors in the development process.
Technology solutions’ development processes are never purely technical, rather they are socio-technical. The problems that AI labs prioritise are those that solve problems they believe are the most pressing, relevant or profitable. Therefore, their lived realities, are very much embedded into the AI systems they build.
Directions for ethnographers
An ethnographer embedded in RD&D can aim to address the following questions (Table 2):
| Theme | Questions |
|---|---|
| Problem definition | What is the process of problem definitions? What are the underlying assumptions of such problem definitions? |
| Data handling practices | What are the data annotating, cleaning, and other handling practices? What are the underlying theoretical frameworks that researchers and developers employed through modelling? |
| Lab-level practices | What is the decision-making process in an AI lab? How is it associated with the broader organizational structure? |
| Lab workers’ practices | Who are the experts, users, or other actors involved in RD&D? What are the social, political, and economic forces that affect the operation of researchers and developers? What is the level of agency of AI labs’ workers in shaping the models’ development? |
(Table 2. Examples of research directions for ethnographers embedded in the AI RD&D.)
The ethnographer’s field site can be within AI-developing companies, academic institutions, or any other entities that might be directly conducting RD&D. The role of the ethnographer can, therefore, be internal and non-adversarial to those entities. Due to the potential constraints that also help make ethnographers non-adversarial, such as security clearance, there exist several levels of access to the specific aspects of an AI lab (Koshiyama et al. 2021; Akula and Garibay 2021). Each level allows access to information of different degrees of sensitivity and privacy, which inevitably shapes the scope of the ethnographic inquiry.
Integrating PD and ethnography can enhance the process and outcomes of AI research. By embedding ethnographers within AI development spaces, organizations can establish immediate feedback loops based on rich qualitative data. This approach resonates with Anderson’s (1994) observation that ethnography serves an analytical function in design, helping practitioners question their assumptions and conventional problem-solving frameworks, and productively inform the design sensibilities. Furthermore, ethnographic documentation captures invaluable meta-knowledge about the development process, including failed approaches that often go unrecorded in formal publications. This comprehensive documentation of both successes and failures can accelerate technological progress, particularly in AI safety research, where understanding ineffective safety measures is as crucial as knowing what works. This methodology encourages a higher reflexivity throughout the development process, ensuring that AI systems are designed with fuller awareness of their potential societal implications and limitations.
3.2. Testing and evaluation (T&E)
About this mega-process
After the development of a prototype, developers carry out or invite third-party entities to test and evaluate their model for a range of purposes: ensuring performance, meeting regulatory requirements, benchmarking for AI safety or responsible development, increasing public trust, etc. For example, ethnographic approaches can be leveraged within pilot tests of AI tools during the pre-deployment stages. This approach would ideally incorporate user profiles to understand AI affordances, which are the perceived and actual capabilities of AI systems that shape how users interact with and employ them. By studying how user groups interpret and utilize these affordances in real-world contexts, developers can better align AI capabilities with user needs and expectations while identifying potential limitations and unintended consequences before wider deployment. This move from development to T&E can be iterative; the results of tests and evaluations may necessitate crucial changes to the models.
Often related to this mega-process is the field of AI auditing, which provides “a structured process whereby an entity’s present or past behaviour is assessed for consistency with relevant principles or norms” to ensure trustworthy AI development (Brundage et al., 2020). AI audits can be conducted internally (i.e. by someone within the AI development entity) or externally (by a third-party entity). While internal auditors bring intimate knowledge of an AI system’s architecture and external auditors provide independent oversight, ethnographers can bridge these perspectives by documenting the lived experience of AI implementation. Combining ethnography with technical AI auditing practices can help ensure that AI systems not only meet formal requirements but also serve their intended purposes effectively within their social contexts.
Ethnographer’s lens
(Figure 3. Mega-process of T&E.)
Observables
In this mega-process, there are two observables an ethnographer can focus on, one straightforward and the other more meta-analytical. The more obvious observable relates to the targets of governance-layer audits mentioned in Mökander et al. (2023). In governance audits, an auditor examines how the AI development entity carries out quality management and control, handles risk management, establishes organisational accountability and incentive structures, and implements testing and verification procedures. In other words, a governance auditor’s observables are process-oriented, which ethnographers will be more than equipped to capture through ongoing studies. As a result, many AI safety evaluations are highly disconnected from empirical evidence and misrepresented as enhancing AI safety due to their confusing, tight connection with capability advancements (Ren et al. 2024). Therefore, there is a need to better cement AI evaluations within empirical evidence and working with users.
The more meta-analytical observable are the various iterations an AI system goes through. Ethnographers can focus here on the iterations of problem definitions after testing and evaluations. The observable may include the developers’ reactions and interpretations of the testing and evaluation results, their brainstorming process to identify solutions or the assumptions and values hidden behind the new design iteration.
Directions for ethnographers
If the digital ethnographer embedded in T&E focuses on the governance structure observable, they can aim to answer the following questions:
| Theme | Questions |
|---|---|
| T&E procedures | What procedures have the AI development entity set up for quality control, risk management, internal reporting, and model testing? How is the implementation of such procedures, and are there exogenous factors (e.g. market demands, press release dates, public relations incidents, regulatory fines) that might affect it? |
| Dynamics within T&E entities | What is the sociocultural dynamic and political hierarchy within the entity? How does the entity establish internal accountability and incentive structures? |
(Table 3. Examples of research directions for ethnographers embedded in the AI T&E (governance structure).)
In the case of examining iteration observable, an ethnographer can address a very similar set of mentioned questions in the RD&D mega-process, but now with a focus on changes throughout iterations:
| Theme | Questions |
|---|---|
| Changes throughout iterations | What are the new problem definitions after testing and evaluations? How are problems identified or reframed? Throughout iterations, what are some changed assumptions or theoretical foundations about the model? Throughout iterations, what are some unchanging assumptions or theoretical foundations about the model? Throughout iterations, what are the priorities of developers in redesigning? What are the persisting features and discarded ones, and how do developers make the distinction? |
| T&E contexts | What are the social, political, and economic forces that affect the priorities of developers? Who are the experts, users, or other actors involved in T&E? |
(Table 4. Examples of research directions for ethnographers embedded in the AI T&E.)
A third-party auditor examining the governance structure observable may begin with key stakeholder interviews and document analysis. It may involve the more intrusive participation observation methods, given the degree of access the AI development entity grants (or is mandated by regulators to grant) to the auditor (Mökander and Floridi 2023).
3.3. Deployment
About this mega-process
As an AI system is deployed, it begins to interact with multiple actors who may employ and incorporate it into their workflow and lifestyle in various ways. As such, AI systems may become embedded in socioeconomic and political environments distinctly different from where they were first developed.
For example, a social media algorithm that prioritizes clicks and engagement rate during violent military conflicts may recommend posts that encourage polarizing languages that undermine peace processes — though this may not have been the intention of the AI labs’ design and development process. Another example is the surge pricing algorithm used by ride-hailing apps (like Uber or Lyft), which aims to balance market demand and supply, and may hinder the mobility of marginalized groups in highly racialized and socioeconomically segregated neighbourhoods (Galdon-Clavell, Nalbantova, and Cutillas 2023).
Ethnographer’s lens
(Figure 4. Mega-process of deployment.)
Observables
During this mega-process, there are multiple observables for an ethnographer to explore. This is also the most studied mega-process as compared to the other two, given that the most popular angle for understanding participation in AI development and AI’s societal impacts is to look at the various use cases.
First, an ethnographer can focus on the heterogeneity of user groups, portraying how local sociocultural norms, value systems, political dynamics, and economic conditions affect users’ interactions with AI systems — thereby moving away from the understanding that users are ‘universal’ and homogeneous. One could highlight the unintended or an extension of expected use cases, cross-cultural and cross-society comparisons of AI use, and the direct and indirect impacts an AI system can have on individual users or groups of individuals. This observable will be especially useful for highlighting the lived experiences of marginalized groups.
Second, the ethnographer can focus on the transfer of context that comes with deployment. When AI systems leave the lab conditions under which they were developed and tested, they face new environments where they are not trained to operate. Digital ethnographers can detect AI systems’ emergent behavioural patterns, reconceptualize their capabilities, and bridge the gap between lab-level evaluations and tests of AI systems and the latter’s real-world performance.
Directions for ethnographers
The value of ethnographers studying the AI lifecycle lies in their ability to go beyond the superficial layer of the mega-processes, by asking more contextual questions and pushing for more reflection. Examples of some of the questions ethnographers could ask include:
| Theme | Questions |
|---|---|
| Understanding the users | Who are the users in question for a particular AI system deployment case? What are the sociocultural norms or ethical values relevant to the users? |
| Use-case contexts | How do local sociopolitical dynamics and economic conditions affect the use case? In which ways do different groups of users employ the AI system? What capabilities do the AI systems exhibit when interacting with a given group of users? Were those capabilities tested at the lab, or were they co-emergent given specific modes of user interaction? What are the differences in contexts, pre- and post-deployment, for the AI system? Do the risk scores given out at lab-level tests and evaluations reflect how the AI system creates societal impacts on the ground? |
(Table 5. Examples of research directions for ethnographers embedded in the AI deployment.)
Another set of questions an ethnographer can answer concerns the meta-analysis that cross-over to other mega-processes:
| Theme | Questions |
|---|---|
| Feedback loop: deployment → RD&D | What do the concrete use cases tell us about the RD&D for AI systems? Are there any systematic issues or patterns in the ethos of AI RD&D to be addressed? How do we create a feedback loop where user needs and experiences may be reflected in the RD&D process? |
| Feedback loop: deployment → T&E | What do the co-emergent capabilities of AI systems through human interactions inform us about the T&E mega-process? Are there any new risk areas that need to be accounted for during RD&D or T&E? Are there any updates that should be made about risk scoring in T&E? Should there be updates made to the methodology of T&E? |
(Table 6. Examples of research directions for ethnographers embedded in the AI deployment.)
The ethnographer’s sites will be where the identified users are. They could carry out inquiries in particularly risky areas (e.g. where disruptions caused by the deployment of a given AI system are expected) or with marginalized or disadvantaged communities in different societies. The ethnographers can be external or internal to AI development entities, but they should be involved in a tight feedback loop where their research findings could directly or indirectly affect RD&D and T&E. For example, one can imagine a multistakeholder collaboration among regulators and AI development entities where digital ethnographers take the lead in investigating particular use cases of AI system deployment in risky contexts.
4. Further considerations
This paper adds to the ongoing discussions on participation in AI development. The outlined framework presents the how-to for embedding ethnographers in the AI lifecycle to help understand the diversity and impact of human practices in AI DDD. This section brings to the forefront some high-level questions regarding the tensions, challenges and limitations of this proposed framework.
Ethnographers can help fill the gap in studying the diversity of forms of participation of various stakeholders in each mega-process and their impact on AI DDD. Adoption of this conceptual framework could offer continuous documentation rather than punctuated studies, bridge the insights from in-the-lab and outside-the-lab field sites and consequently, produce a more holistic view of practices and structures involved in the entire AI lifecycle.
4.1. Tensions
Negotiating field access
In the case of AI labs, negotiating field access may turn out to be especially difficult given the low transparency levels across AI development spaces and the cultural barriers to becoming a legitimate insider. Firstly, AI labs possess proprietary and sensitive information accessible to only a handful of people--or the insiders. In this context, it is challenging for an ethnographer tasked with studying the lab’s practices to gain the potential insiders’ trust. A form of reciprocity can be helpful where the ethnographer studies questions that are relevant to the developers and designers. However, balancing the protection of AI labs’ interests and an inductive research agenda may turn out to be challenging. Secondly, different communities embody specific habitus, values and views that may act as gatekeepers to accessing the people and the information inside them (Bourdieu 1977). Ethnographers may struggle to be seen as part of the AI lab community should they not speak the same technical language, learn the terms of references, or share similar skills and practices (Wu 2024). As a consequence, insiders may withhold information or experience observation bias, where they unintentionally interpret or selectively perceive data to confirm their preexisting beliefs, rather than objectively documenting what is actually happening.
Scalable Oversight versus Observation
Scalable oversight aims to supervise AI systems on a mass scale (Bowman et al. 2022; Vesely and Kim 2024). As more and more AI systems permeate each industry and across geographic boundaries, the challenge lies in scaling such oversight to systems operating at global levels and in sensitive domains (e.g. healthcare) (van Voorst and Ahlin 2024). Observing human interactions with AI systems at scale highlights the tension between the need for centralised oversight and the reality of localised and distributed human behaviours. While centralized monitoring of AI systems may achieve consistency, it often overlooks the nuanced dynamics of diverse contexts in which different human users operate. Methods such as ethnography can capture nuanced dynamics; nevertheless, such observation alone is inherently limited in scale, as it focuses on localized contexts and can be labour-intensive.
4.2. Limitations
Ethnographic methods can provide valuable insights and viewpoints into the mechanics of AI DDD, but they require a high level of access given to the ethnographer for an open-ended fieldwork inquiry, trusting relationships and transparency on data and practices (van Voorst and Ahlin 2024), and informed consent of the study participants.
Limitations of participant observation
The adoption of more inclusive and collaborative approaches that favour co-creation and context-sensitive needs is an essential goal. However, these approaches also come with the challenge of navigating researcher-informant power dynamics, especially in end-user contexts where the deployed AI systems may increase the informants’ vulnerability. There is a need for a continuous ethically engaged presence in the field that considers participants’ agency, and representation and mitigates any potential exploitation (Watts 2010).
Limitations of the conceptualisation of the AI lifecycle
The AI lifecycle is far from straightforward. The broader socio-technical dynamics and the roles of various stakeholders involved in AI systems are inherently complex and beyond binary classification. As a result, the mega-process framework used in this white paper, while useful as a conceptual tool for visualising stages in AI development and deployment, is not comprehensive or prescriptive. Rather, it is based on a set of specific assumptions about many factors, such as the identified users, the directions and positions of feedback loops, and power asymmetries in AI systems.
Limitations of reflexivity
Ethnography requires the researcher’s immersion in the field. However, over time, the degree and intensity of embeddedness may impact the ethnographer’s ability to remain reflexive and analytical (Geertz 1977), both about their own positionality and those they study. In the case where the ethnographer simultaneously pursues work for an AI lab in return for field access, their ability to critically study the observed insiders’ practices may be compromised.
Limitations of ethnography
Ultimately, the very characteristics that make the ethnographic research methodology valuable, may turn out to be its limitations. Ethnography facilitates a deep understanding of specific cultures and people’s practices. However, a “complete” analysis of a particular setting is impossible (Shapiro 1994). That is because ethnographers’ practices are essentially interpretive and perspectival, and thus, partial (Suchman 2002). Reproducibility of the study process and generalisation of its findings, in this case, are limited due to the dynamic nature of the study settings and unique qualities or biases embodied by specific researchers.
5. Conclusion
The increasing integration of AI systems into diverse and complex sociocultural environments necessitates a robust and nuanced approach to their DDD. This white paper has argued that while PD methodologies offer valuable insights by prioritizing the needs and perspectives of end users—particularly those historically marginalized, ethnographic approaches can provide a more comprehensive account that examines the entirety of the AI lifecycle. Ethnography offers a unique method to capture the intricate interactions between developers, designers, auditors, and users, revealing how personal worldviews, organizational structures, and cultural environments profoundly influence technological innovation.
Hence, this white paper presents a conceptual framework for how ethnographers can be embedded throughout the AI DDD, broken down into three mega-processes (RD&D, T&E, and deployment). It focuses on the different observables and research directions that ethnographers can take as they unravel the often invisible sociocultural contexts and overlooked human practices that affect AI DDD.
The key contributions of this framework include challenging the often-used assumption of universally homogenous users during design and development, making developers' practices and taken-for-granted assumptions transparent, and creating robust feedback loops between different stages of AI DDD. By bringing these invisible aspects to the forefront, the proposed ethnographic approach offers a more nuanced, context-sensitive method for examining how AI DDD truly unfolds across complex sociocultural contexts.
However, the proposed approach is not without limitations. Negotiating field access, maintaining researcher reflexivity, and scaling ethnographic observations remain significant challenges. The complexity of AI systems and the dynamic nature of technological contexts further complicate comprehensive analysis. Future research should focus on developing practical implementation strategies for this framework, adding more case studies that demonstrate both the value and limitations of the framework, and refining methodologies for embedding ethnographers across diverse AI development environments.
6. References
Akula, Ramya, and Ivan Garibay. 2021. “Audit and Assurance of AI Algorithms: A Framework to Ensure Ethical Algorithmic Practices in Artificial Intelligence.” arXiv. https://doi.org/10.48550/arXiv.2107.14046.
Anderson, R. J. 1994. “Representations and Requirements: The Value of Ethnography in System Design.” Human-Computer Interaction 9 (2): 151–82. https://doi.org/10.1207/s15327051hci0902_1.
Beede, Emma, Elizabeth Baylor, Fred Hersch, Anna Iurchenko, Lauren Wilcox, Paisan Ruamviboonsuk, and Laura M. Vardoulakis. 2020. “A Human-Centered Evaluation of a Deep Learning System Deployed in Clinics for the Detection of Diabetic Retinopathy.” In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems, 1–12. CHI ’20. New York, NY, USA: Association for Computing Machinery. https://doi.org/10.1145/3313831.3376718.
Birhane, Abeba, William Isaac, Vinodkumar Prabhakaran, Mark Diaz, Madeleine Clare Elish, Iason Gabriel, and Shakir Mohamed. 2022. “Power to the People? Opportunities and Challenges for Participatory AI.” In Proceedings of the 2nd ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization, 1–8. EAAMO ’22. New York, NY, USA: Association for Computing Machinery. https://doi.org/10.1145/3551624.3555290.
Blomberg, Jeanette, Mark Burrell, and Greg Guest. 2002. “An Ethnographic Approach to Design.” In , 964–86. https://doi.org/10.1201/b11963-52.
Blomberg, Jeanette, Jean Giacomi, Andrea Mosher, and Pat Swenton-Wall. 1993. “Ethnographic Field Methods and Their Relation to Design.” In Participatory Design. CRC Press.
Blomberg, Jeanette, and Helena Karasti. 2012. “Ethnography: Positioning Ethnography within Participatory Design.” In Routledge International Handbook of Participatory Design, edited by Jesper Simonsen and Toni Robertson, 0 ed., 106–36. Routledge. https://doi.org/10.4324/9780203108543-12.
Bourdieu, Pierre. 1977. Outline of a Theory of Practice. Translated by Richard Nice. Cambridge Studies in Social and Cultural Anthropology. Cambridge: Cambridge University Press. https://doi.org/10.1017/CBO9780511812507.
Bowman, Samuel R., Jeeyoon Hyun, Ethan Perez, Edwin Chen, Craig Pettit, Scott Heiner, Kamilė Lukošiūtė, et al. 2022. “Measuring Progress on Scalable Oversight for Large Language Models.” arXiv. https://doi.org/10.48550/arXiv.2211.03540.
Buyl, Maarten, Alexander Rogiers, Sander Noels, Iris Dominguez-Catena, Edith Heiter, Raphael Romero, Iman Johary, Alexandru-Cristian Mara, Jefrey Lijffijt, and Tijl De Bie. 2024. “Large Language Models Reflect the Ideology of Their Creators.” arXiv. https://doi.org/10.48550/arXiv.2410.18417.
Cao, Yong, Li Zhou, Seolhwa Lee, Laura Cabello, Min Chen, and Daniel Hershcovich. 2023. “Assessing Cross-Cultural Alignment between ChatGPT and Human Societies: An Empirical Study.” arXiv. https://doi.org/10.48550/arXiv.2303.17466.
Crabtree, Andy. 1998. “Ethnography in Participatory Design.” In Proceedings of the 1998 Participatory Design Conference, 93–105. Seattle, Washington, USA: Computer Professionals Social Responsibility.
Cratsley, Maya J., and Nathanael J. Fast. 2024. “‘Inventor’s Bias’ at Work: When Low-Performing Algorithms Seem Fair.” International Journal of Human–Computer Interaction 40 (1): 24–32. https://doi.org/10.1080/10447318.2023.2224954.
Delgado, Fernando, Stephen Yang, Michael Madaio, and Qian Yang. 2023. “The Participatory Turn in AI Design: Theoretical Foundations and the Current State of Practice.” In Proceedings of the 3rd ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization, 1–23. EAAMO ’23. New York, NY, USA: Association for Computing Machinery. https://doi.org/10.1145/3617694.3623261.
Galdon-Clavell, Gemma, Iliyana Nalbantova, and Sergi Cutillas. 2023. “Auditing Uber, Bolt and Cabify: Algorithmic Audit Uncovers Compliance Issues.” Association Eticas Research and Innovation.
Geertz, Clifford. 1977. The Interpretation Of Cultures. Basic Books.
Gill, Karamjit S. 1989. “Reflections on Participatory Design.” AI & SOCIETY 3 (4): 297–314. https://doi.org/10.1007/BF01908620.
Ibrahim, Lujain, Saffron Huang, Lama Ahmad, and Markus Anderljung. 2024. “Beyond Static AI Evaluations: Advancing Human Interaction Evaluations for LLM Harms and Risks.” arXiv. https://doi.org/10.48550/arXiv.2405.10632.
Koshiyama, Adriano, Emre Kazim, Philip Treleaven, Pete Rai, Lukasz Szpruch, Giles Pavey, Ghazi Ahamat, et al. 2021. “Towards Algorithm Auditing: A Survey on Managing Legal, Ethical and Technological Risks of AI, ML and Associated Algorithms.” SSRN Scholarly Paper. Rochester, NY. https://doi.org/10.2139/ssrn.3778998.
Kuthiala, Nitya, and Keegan McBride. 2025. “How Silicon Valley Developed Digital Dating Platforms Are Transforming Love and Relationship Culture in India.” SSRN Scholarly Paper. Rochester, NY: Social Science Research Network. https://doi.org/10.2139/ssrn.5097768.
Liu, Chen, Fajri Koto, Timothy Baldwin, and Iryna Gurevych. 2024. “Are Multilingual LLMs Culturally-Diverse Reasoners? An Investigation into Multicultural Proverbs and Sayings.” In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), edited by Kevin Duh, Helena Gomez, and Steven Bethard, 2016–39. Mexico City, Mexico: Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.naacl-long.112.
Masoud, Reem I., Ziquan Liu, Martin Ferianc, Philip Treleaven, and Miguel Rodrigues. 2024. “Cultural Alignment in Large Language Models: An Explanatory Analysis Based on Hofstede’s Cultural Dimensions.” arXiv. https://doi.org/10.48550/arXiv.2309.12342.
Mökander, Jakob, and Luciano Floridi. 2023. “Operationalising AI Governance through Ethics-Based Auditing: An Industry Case Study.” AI and Ethics 3 (2): 451–68. https://doi.org/10.1007/s43681-022-00171-7.
Mökander, Jakob, Jonas Schuett, Hannah Rose Kirk, and Luciano Floridi. 2023. “Auditing Large Language Models: A Three-Layered Approach.” SSRN Scholarly Paper. Rochester, NY. https://doi.org/10.2139/ssrn.4361607.
Oh, Changhoon, Seonghyeon Kim, Jinhan Choi, Jinsu Eun, Soomin Kim, Juho Kim, Joonhwan Lee, and Bongwon Suh. 2020. “Understanding How People Reason about Aesthetic Evaluations of Artificial Intelligence.” In Proceedings of the 2020 ACM Designing Interactive Systems Conference, 1169–81. DIS ’20. New York, NY, USA: Association for Computing Machinery. https://doi.org/10.1145/3357236.3395430.
Ren, Richard, Steven Basart, Adam Khoja, Alice Gatti, Long Phan, Xuwang Yin, Mantas Mazeika, et al. 2024. “Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?” arXiv. https://doi.org/10.48550/arXiv.2407.21792.
Seaver, Nick. 2013. “Knowing Algorithms.” In digitalSTS: A Field Guide for Science & Technology Studies. Princeton, New Jersey, United States: Princeton University Press. https://digitalsts.net/essays/knowing-algorithms/.
———. 2017. “Algorithms as Culture: Some Tactics for the Ethnography of Algorithmic Systems.” Big Data & Society 4 (2): 2053951717738104. https://doi.org/10.1177/2053951717738104.
Shapiro, Dan. 1994. “The Limits of Ethnography: Combining Social Sciences for CSCW.” In Proceedings of the 1994 ACM Conference on Computer Supported Cooperative Work, 417–28. CSCW ’94. New York, NY, USA: Association for Computing Machinery. https://doi.org/10.1145/192844.193064.
Suchman, Lucy A. 2002. “Practice-Based Design of Information Systems: Notes from the Hyperdeveloped World.” The Information Society 18 (2): 139–44. https://doi.org/10.1080/01972240290075066.
Vesely, Stepan, and Byungdoo Kim. 2024. “Survey Evidence on Public Support for AI Safety Oversight.” Scientific Reports 14 (1): 31491. https://doi.org/10.1038/s41598-024-82977-5.
Voorst, Roanne van, and Tanja Ahlin. 2024. “Key Points for an Ethnography of AI: An Approach towards Crucial Data.” Humanities and Social Sciences Communications 11 (1): 1–5. https://doi.org/10.1057/s41599-024-02854-4.
Watts, Jacqueline H. 2010. “Ethical and Practical Challenges of Participant Observation in Sensitive Health Research.” International Journal of Social Research Methodology 14 (4): 301–12.
Wu, Yung-Hsuan. 2024. “Capturing the Unobservable in AI Development: Proposal to Account for AI Developer Practices with Ethnographic Audit Trails (EATs).” AI and Ethics, September. https://doi.org/10.1007/s43681-024-00535-1.
Zytko, Douglas, Pamela J. Wisniewski, Shion Guha, Eric P. S. Baumer, and Min Kyung Lee. 2022. “Participatory Design of AI Systems: Opportunities and Challenges Across Diverse Users, Relationships, and Application Domains.” In Extended Abstracts of the 2022 CHI Conference on Human Factors in Computing Systems, 1–4. CHI EA ’22. New York, NY, USA: Association for Computing Machinery. https://doi.org/10.1145/3491101.3516506.
Footnotes
-
Blomberg et al., 2002:966. Xerox Palo Alto Research Center’s (PARC’s) motto; the center worked on intergrating ethnographic studies in innovation design. ↩
-
In this paper, the term “AI system” refers to an algorithmic model that a computer builds partially without human intervention after observing some data and recognizing patterns from such data. This paper does not make the distinction between an AI model and an AI system for simplicity. ↩
-
In this paper, AI laboratories (labs) refer to any research, design, and developmental spaces of AI systems. It can be an entity by itself, or it can also be nested within a bigger institution. ↩
-
This paper focuses on the simplest and most common profiles of actors featured within AI DDD and is by no means comprehensive or exhaustive. The actor list that can be involved and, therefore, studied in ethnography can be further expanded and is dependent on the context of each ethnography. ↩
-
As with Figure 3. and Figure 4., the orange circles indicate the lens of an ethnographer in that mega-process. ↩




