Title: “Bettering Data”: The Role of Everyday Language and Visualization in Critical Novice Data Work
I nformed by criti cal data literacy efforts to promote social justice, this paper uses qualitative m ethods and data collected during two years of w orkplace ethnograph y t o c haracterize the notion of critical n ovice data work . Specifically, we analyze everyday language used by novice data workers at DataWorks, an organization that trains and employs historically excluded populations to work with community data sets. We also characterize challenges faced by these workers in b oth cleaning and being critical of data d uring a pr oj ect f ocused o n p olice -c ommunity re l ations. Finally, we highlight novel approaches to visualizing data that the workers developed during this project, derived from data cleaning and everyday experience. Findings and discussion highlight the generative power of everyday language and visualization for critical novice data work, as well as challenges and opportunities to foster critical data literacy with novice data workers in the workplace.  more » « less
Award ID(s):
1951818
PAR ID:
10662779
Author(s) / Creator(s):
; ; ; ; ; ;
Publisher / Repository:
JSTOR
Date Published:
Journal Name:
Educational technology society
ISSN:
1436-4522
Format(s):
Medium: X
Sponsoring Org:
National Science Foundation
More Like this
  1. null (Ed.)
    In this paper, we describe and analyze a workshop developed for a work training program called DataWorks. In this workshop, data workers chose a topic of their interest, sourced and processed data on that topic, and used that data to create presentations. Drawing from discourses of data literacy; epistemic agency and lived experience; and critical race theory, we analyze the workshops’ activities and outcomes. Through this analysis, three themes emerge: the tensions between epistemic agency and the context of work, encountering the ordinariness of racism through data work, and understanding the personal as communal and intersectional. Finally, critical race theory also prompts us to consider the very notions of data literacy that undergird our workshop activities. From this analysis, we ofer a series of suggestions for approaching designing data literacy activities, taking into account critical race theory. 
    more » « less
  2. In this paper, we describe and reflect upon the development of critical consciousness and workplace democracy within an experimental workplace called DataWorks. Through DataWorks, we hire adults from communities historically minoritized in computing education and data careers, and train them in entry-level data skills developed through work on client projects. In this process, workers gain a range of skills. Some of these skills are technical, such as programming for data analysis; some are managerial, such as scoping and bidding projects; others are social, perhaps even political, such as the ability to say No to projects. In what follows, we describe a workshop series developed to build the workers' critical literacy and consciousness about their data work, specifically regarding the use of data in machine learning systems. After that, we describe a data project the workers questioned and resisted because they determined the work to be harmful. In that process, they demonstrated and enacted a critical consciousness towards data and machine learning. Reflecting on this enactment of data-focused critical consciousness, we identify themes that characterize a democratic workplace, describe the work of designing for organizational action and institutional relations, and discuss how worker and researcher positionality affects this work. In doing so, we argue for enabling workers to resist and refuse harmful data work and challenge the standard power structures of academic research and data work. 
    more » « less
  3. Many AI system designers grapple with how best to collect human input for different types of training data. Online crowds provide a cheap on-demand source of intelligence, but they often lack the expertise required in many domains. Experts offer tacit knowledge and more nuanced input, but they are harder to recruit. To explore this trade off, we compared novices and experts in terms of performance and perceptions on human intelligence tasks in the context of designing a text-based conversational agent. We developed a preliminary chatbot that simulates conversations with someone seeking mental health advice to help educate volunteer listeners at 7cups.com. We then recruited experienced listeners (domain experts) and MTurk novice workers (crowd workers) to conduct tasks to improve the chatbot with different levels of complexity. Novice crowds perform comparably to experts on tasks that only require natural language understanding, such as correcting how the system classifies a user statement. For more generative tasks, like creating new lines of chatbot dialogue, the experts demonstrated higher quality, novelty, and emotion. We also uncovered a motivational gap: crowd workers enjoyed the interactive tasks, while experts found the work to be tedious and repetitive. We offer design considerations for allocating crowd workers and experts on input tasks for AI systems, and for better motivating experts to participate in low-level data work for AI. 
    more » « less
  4. In this paper, we propose situating data literacy within a humanities framework as a complementary method in data literacy education, using network visualization as a tool. As data becomes increasingly essential in understanding our world, there is a growing need for a multidisciplinary approach to data literacy. We argue that network visualization techniques can foster deeper engagement with data's complexities in certain contexts. We explore various definitions and frameworks of data literacy, emphasizing context, critical thinking, and holistic understanding, and highlight network visualization as a tool to explore data within a humanistic framework. By using networks, students can develop a critical and contextual understanding of data, preparing them to navigate the complexities of our data-driven society more effectively. 
    more » « less
  5. Many publicly available datasets exist that can provide factual answers to a wide range of questions that benefit the public. Indeed, datasets created by governmental and nongovernmental organizations often have a mandate to share data with the public. However, these datasets are often underutilized by knowledge workers due to the cumbersome amount of expertise and embedded implicit information needed for everyday users to access, analyze, and utilize their information. To seek solutions to this problem, this paper discusses the design of an automated process for generating questions that provide insight into a dataset. Given a relational dataset, our prototype system architecture follows a five-step process from data extraction, cleaning, pre-processing, entity recognition using deep learning, and questions formulation. Through examples of our results, we show that the questions generated by our approach are similar and, in some cases, more accurate than the ones generated by an AI engine like ChatGPT, whose question outputs while more fluent, are often not true to the facts represented in the original data. We discuss key limitations of our approach and the work to be done to bring to life a fully generalized pipeline that can take any data set and automatically provide the user with factual questions that the data can answer. 
    more » « less