<?xml-model href='http://www.tei-c.org/release/xml/tei/custom/schema/relaxng/tei_all.rng' schematypens='http://relaxng.org/ns/structure/1.0'?><TEI xmlns="http://www.tei-c.org/ns/1.0">
	<teiHeader>
		<fileDesc>
			<titleStmt><title level='a'>Measuring Temporal Awareness for Human-Aware AI</title></titleStmt>
			<publicationStmt>
				<publisher>Sage Journals</publisher>
				<date>09/01/2023</date>
			</publicationStmt>
			<sourceDesc>
				<bibl> 
					<idno type="par_id">10506881</idno>
					<idno type="doi">10.1177/21695067231192635</idno>
					<title level='j'>Proceedings of the Human Factors and Ergonomics Society Annual Meeting</title>
<idno>1071-1813</idno>
<biblScope unit="volume">67</biblScope>
<biblScope unit="issue">1</biblScope>					

					<author>Margaret A. Gray</author><author>Zhuorui Yong</author><author>Abhijan Wasti</author><author>Esa M. Rantanen</author><author>Jamison R. Heard</author>
				</bibl>
			</sourceDesc>
		</fileDesc>
		<profileDesc>
			<abstract><ab><![CDATA[This research investigated human performance in response to task demands that may be used to convey information about the human to an artificial agent. We performed an experiment with a dynamic time-sharing task to investigate participants development of temporal awareness of the task event unfolding in time. Temporal awareness as an extension, or a special case, of situation awareness, may provide for useful measures of covert mental models applicable to numerous tasks and for input to human-aware AI agents. Temporal awareness measures may be used to classify human performance into the control modes in the contextual control model (COCOM): scrambled, opportunistic, tactical, and strategic. Twenty-one participants participated in a within subjects experiment with an abstract task of resetting four independent timers within their respective windows of opportunity. The results show that temporal measures of task performance are sensitive to changes in task disruptions and difficulty and therefore have promise for human-aware AI.]]></ab></abstract>
		</profileDesc>
	</teiHeader>
	<text><body xmlns="http://www.tei-c.org/ns/1.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink">
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Introduction</head><p>Human-aware artificial intelligence (AI) refers to autonomous systems that can effectively interact, collaborate, and team with humans for a variety of tasks <ref type="bibr">(Kambhampati, 2020)</ref>. Human-robot collaboration is a primary area for development of human-aware robots <ref type="bibr">(Kumar, Arora, &amp; Sahin, 2019;</ref><ref type="bibr">Singh &amp; Heard, 2022)</ref>, but to make any kind of AI-driven automation truly human-aware, there must be a way for the automated agents to sense various aspects of human behavior and performance to adapt accordingly and to offer a truly collaborative experience to the human. Human sensing has a long history, spanning the dead man's switches in electric streetcars and subway trains from the last century to increasingly sophisticated driver monitoring systems in highly automated automobiles in the present day <ref type="bibr">(Hecht et al., 2019)</ref>. However, it may be argued that human sensing research is lagging behind accelerating development of machine learning (ML) and AI-driven automation applications in all areas of life, not only in human-robot collaboration in industrial settings or in self-driving cars.</p><p>Several criteria may be developed for human sensing systems. They should be unobtrusive so as not to interfere with the human task performance in any way. They should not rely on wearable instrumentation requiring lengthy set-up or restricting human motions in task performance. They should be applicable to a wide variety of human-automation interactions (HAI) <ref type="bibr">(Kaber, 2018)</ref> or human-autonomy teaming (HAT) <ref type="bibr">(O'Neill, McNeese, Barron, &amp; Schelble, 2022)</ref>. Finally, they should account for a wide variety of cognitive styles, strategies, and individual differences in humans <ref type="bibr">(Feigh, 2011)</ref>.</p><p>Ubiquitous and multiple-dynamic <ref type="bibr">(Reason, 1990)</ref> interactions between two fundamentally different agents, humans as analog beings <ref type="bibr">(Norman, 1998)</ref> and digital computers, each relying on imperfect and different but interactively and dynamically shaped models of each other, present additional research problems <ref type="bibr">(Begerowski, Hedrick, Waldherr, Mears, &amp; Shuffler, 2023;</ref><ref type="bibr">O'Neill, Flathmann, McNeese, &amp; Salas, 2023;</ref><ref type="bibr">O'Neill et al., 2022;</ref><ref type="bibr">Stowers, Brady, MacLellan, Wohleber, &amp; Salas, 2021)</ref>. AI is trained by experience with human interactions, but these interactions are also influenced by the human experience with the AI agent. AI systems model humans as biological neural nets being trained by the systems themselves <ref type="bibr">(Christian, 2020)</ref>. Models and methods to enable the integration of humans and technologies in ways that optimally utilize their key strengths are needed as well <ref type="bibr">(Hagenow et al., 2021a</ref><ref type="bibr">(Hagenow et al., , 2021b;;</ref><ref type="bibr">Pearce, Mutlu, Shah, &amp; Radwin, 2018;</ref><ref type="bibr">Schoen, Henrichs, Strohkirch, &amp; Mutlu, 2020)</ref>.</p><p>In this paper we describe an experimental paradigm to investigate human performance and behavioral indices in dynamic tasks. Our research is based on two theoretical frameworks, situation awareness (SA) <ref type="bibr">(Endsley, 1988)</ref> and the contectual control model (COCOM) <ref type="bibr">(Hollnagel, 1993)</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Temporal Awareness</head><p>Dynamic systems refer to environments where one must keep track of and respond to multiple changing variables. Successful control of dynamic systems implies that the users have a "mental model" of the system, allowing the user to predict system behavior and the consequences of their inputs to it. Time is an integral dimension of these dynamic systems, and is an inherent component and constraint in nearly every human activity. Having good awareness of events unfolding in time, or good temporal awareness, is crucial in creating an effective mental models of dynamic systems.</p><p>Situation awareness (SA) represents a specific, dynamic, aspect of mental models. The nearly universally accepted definition of SA by <ref type="bibr">Endsley (1988)</ref> involves three levels: (1) perception of the elements in the environment, (2) comprehension of their meaning, and (3) their projection into the future. <ref type="bibr">Endsley (2000)</ref> also made a distinction between mental models as representative of static knowledge about a system, whereas SA embodies a situation model, which is an extraction of time-and event-specific information from the underlying mental model. As a theoretical construct, SA has proved to be somewhat elusive, defying attempts to postulate plausible mechanisms behind it, and even its quantification in various settings. Time as a variable common to systems' dynamics and human performance may be used as means to quantify SA.</p><p>Temporal awareness as an extension, or a special case of SA, may provide for useful measures of covert mental models applicable to numerous tasks. Appropriate task prioritization is a key performance metric in many tasks. Task prioritization further depends on accurate estimation of three temporal task parameters: (1) the time when the task becomes "available", or the time when a window of opportunity (WO) to perform it opens, (2) the latest time by which the task must be completed, or the closing of the WO, and (3) the time required to perform the task <ref type="bibr">(Rantanen, 2009)</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Contextual Control Model</head><p>The COCOM model developed by <ref type="bibr">Hollnagel (1993</ref><ref type="bibr">Hollnagel ( , 1998) )</ref> identified several parameters that may yield useful and practical measures of operator performance. This model distinguishes four control modes: scrambled, opportunistic, tactical, and strategic. In the scrambled mode, human performance is haphazard and unpredictable, without planning, and can be best described as a state of momentary panic, representing a complete loss of SA. The opportunistic mode is only slightly better in terms of performance or SA; the operator merely responds to the most salient events (e.g., alarms) but is not able to plan actions or predict their consequences. The tactical control mode involves planning and the operator is in control of the situation or the system, implying a moderately good SA. Finally, in strategic control mode the operator is in complete control of the task, able to consider the global context, and exhibiting good SA. Human performance in the first two modes may be characterized as reactive and in the latter two modes as proactive. Reactive and proactive behavior may be distinguishable in the timing of actions, offering a potential means for performance measurement.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>The Experimental Paradigm</head><p>The experimental task in this research is an abstract timesharing task, originally developed by Rantanen <ref type="bibr">(Levinthal &amp; Rantanen, 2004;</ref><ref type="bibr">Rantanen &amp; Levinthal, 2005)</ref> to study workload in dynamic task settings and used in other research since <ref type="bibr">(Kulom&#228;ki, Oksama, Rantanen, &amp; Hy&#246;n&#228;, 2022)</ref>. The task is performance-dependent; speed and accuracy of performance in one trial affect the onset of subsequent trials. In this way, the task mimics the dynamic nature of real-world scenarios. The measures derived from participant responses to the task demands have been shown to be sensitive to time pressure, measured as the ratio of time required to perform a task to time available to do so <ref type="bibr">(Levinthal &amp; Rantanen, 2004)</ref> and applicable to other tasks as well <ref type="bibr">(Rantanen, 2009;</ref><ref type="bibr">Rantanen &amp; Levinthal, 2005)</ref>. The current study was the first in a planned series of experiments to further develop and test human measures in a variety of tasks that could be used as input in human-aware AI and adaptive automation applications.</p><p>From a task timeline, several measures of temporal awareness may be derived, such as the proper prioritization of tasks and the "timeliness" of performance. In particular, it may be possible to measure the elapsed time from opening of a window of opportunity on individual tasks to an observable action on that task; good temporal awareness is manifested in timely performance on tasks, or consistently short "time to first action" from the opening of the window. Degradation of temporal awareness in turn is manifested in increasing variability in attending to tasks and late performance (completion of tasks after closing of the window of opportunity).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Purpose of the Research and Hypotheses</head><p>The purpose of this study was to investigate potential measures for creating representations of human temporal awareness that may provide useful feedback to humanautonomy teaming systems. The following hypotheses were developed: H1 In an "easy" condition, participants should learn the regular pattern, or sequence, of several substasks, having good performance indicated by little variation in performing the tasks relative to their respective WOs. H2 As the condition changes (surprising the participants) and the (hypothetically) learned subtask performance pattern is disrupted, the participants' performance should become poor, indicated by increased variability in performing the subtasks relative to their respective WOs. H3 The participants will learn the new pattern in time, again exhibiting little variation in performing the subtasks relative to their respective WOs.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Method Participants</head><p>Twenty-one students recruited from an undergraduate psychology course volunteered to participate in exchange for course credit. This research was approved by the Office of Human Subject Research at the Rochester Institute of Technology, and all participants gave their informed consent to participate. The participants were 18 to 21 years of age (M = 18.8 years); 11 self-identified as female, 8 as male, and 2 as non-binary. The participants had different ethnicities including Hispanic, White, Black, and Asian. All participants had the highest completed education of high school or equivalent. Twenty participants were native English speakers; one participant was deaf, and one hard of hearing. All but six participants were avid players of video games.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Apparatus</head><p>A software program, developed on PsychoPy and Python3, simulated a time-sharing task. A computer screen was divided into four panes, which were masked. To unmask a pane and reveal a subtask in it, the participant had to move a cursor to that pane. Once the cursor was moved to another pane, the previous pane was again masked and a different subtask revealed in the pane where the cursor was. The subtasks consisted of progress bars, which were to be reset by typing a 4-digit code within their respective WOs, indicated on the bars. Moreover, when the WO opened the bar turned green and When the WO closed without reset code entered the bar turned red. The participants' task was to monitor the status of four independent progress bars and reset them within their respective WOs (i.e., before the WO closed; a bar could, and should, be reset late, or after closing the WO). Figure <ref type="figure">1</ref> shows the experimental task.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Independent Variables</head><p>We designed two different experimental conditions. The first three minutes in the task were considered "easy", with a relatively leisurely pace at which the bars were to be reset to allow participants develop a good temporal awareness of the subtasks (bar resets) to be performed. After three minutes the bar speed was increased and the sequence in which the bars were to be reset was changed for the second half, or another 3 minutes, of the experiment, surprising the participants. Re-learning a new sequence of subtasks and the overall faster pace would make re-acquiring temporal awareness difficult, hypothetically reflected in observable performance. Specific task parameters are presented in Table <ref type="table">1</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Dependent Variables</head><p>The experimental program recorded several time-stamped events and the participants' actions. We also tracked the mouse position and keyboard events by the program. Table <ref type="table">2</ref> shows the raw measures, and Table <ref type="table">3</ref> shows the key events.</p><p>From these data we derived the time to first action (TFA) within each pane, calculated from the length of time between the opening of the WO to when the user typed in the first digit in the reset code. This was the primary performance metric representing the timeliness of performance in the four subtasks.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Experimental Design</head><p>This was a within-subject design. We divided the duration of the experiment into six 1-minute time epochs, the first three representing the "Easy" condition and the last three representing the "Hard" condition. We wanted to examine learning of the task during the first half of the experiment and the effect of quickening the overall task pace and disruption of learned patterns on task performance.  The WO opened within an individual pane 4.WO-Close</p><p>The WO closed within an individual pane 5. Pane reset User entered the correct input, the progress bar goes back to 0, and a new question is assigned</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Procedure</head><p>The researchers briefly introduced the study to the participants. First, participants completed a demographic information survey, followed by the researchers explaining the task and how to interact with the program. Participants read, agreed, and signed the consent form and started the experiment when ready. The experiment lasted a total of 6 minutes. After the experiment program was finished, researchers conducted a brief one-on-one interview with participants to identify what changes they noticed, their strategies to handle the task, and how they evaluated their own performance.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Results</head><p>Each minute within the six-minute experimental trial was split up into 1-minute "epochs", with the "Easy" condition being divided into epochs 1-3, and the "Hard" condition into epochs 4-6. For each epoch, the TFA values were not normally distributed, as is common with timing data, and as shown in Figure <ref type="figure">2</ref>. Log transformations worked well to normalize the data. Figure <ref type="figure">3</ref> shows the mean and 95% confidence interval for TFA for each of the six epochs (backtransformed from the log scale), and Table <ref type="table">4</ref> lists the values of all descriptive statistics on both the log and backtransformed scales. Within the "Easy" condition of the experiment, the mean log TFA decreased from 3.22 to 2.44 to 2.34, and the standard deviation of log TFA also decreased from 1.08 to 0.99 to 0.88. Within the "Hard" condition, the mean log TFA increased from 1.24 to 1.31 from epochs 4 to 5, and then in epoch 6 decreased to 1.26. The standard deviation of log TFA decreased from 0.80 to 0.78 to 0.73.</p><p>Between the "Easy" and "Hard" conditions, from epochs 3 to 4, the mean log TFA increased by 0.49, from 0.85 to 1.24. The standard deviation of the log TFA decreased from 0.88 to 0.80. A two-sample Wilcoxon rank test (Mann-Whitney U test) was performed to compare the two conditions (epochs 1-3 vs epochs 4-6), and produced a p-value of less than 0.0001.</p><p>These trends support our hypotheses. Within the first three minutes (epochs) and in the "Easy" condition, participants clearly learned the task and improved in their performance, resetting the bars soon after opening the WO (decreasing mean log TFA) and exhibiting increasing consistency in their performance (decreasing standard deviation of log TFA). Interruption of this good performance by increasing the task pace and dis-rupting the learned reset pattern resulted in poorer performance in epoch 4 (lower mean log TFA), but decreased variability (decrease in standard deviation of log TFA). Although the performance appears to have been stabilized in epochs 5 and 6, it never reached the same low level achieved in epoch 3 within the remaining experimental duration.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Discussion</head><p>Within the ''Easy" condition, the decrease in mean log TFA and the decrease in standard deviation of log TFA provides Measurable Description</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Mouse position</head><p>The current X and Y coordinates of the cursor on the screen, between -1.000 and 1.000 Bar progress</p><p>For each pane, the current position of the progress bar, recorded as a number between 0.00 and 100.00 Pane keystrokes For each pane, the current keystrokes that have been entered Mouse-in-pane Which pane the cursor is currently in Event-in-pane Which pane the event occurred in evidence to support our H1 hypothesis that participants will improve their performance condition as they spend time getting used to the pattern. Within the "Hard" condition, it is interesting that there does not seem to be much of a significant decline in mean log TFA or a decrease in the standard deviation of mean log TFA. We had originally hypothesized (H3) that, as participants learned the new pattern, they would display a similar trend of a decreasing mean log TFA and decreasing standard deviation of log TFA. Our findings do not support that hypothesis.</p><p>Between the "Easy" and "Hard" conditions, there was a jump in TFA, suggesting that the change in conditions did impact performance. A Mann-Whitney U-test showed a significant difference between the two conditions (p &lt; 0.0001). We hypothesized (H2) that, as the condition changes, the reset pattern is disrupted, and variability in the participants' performance would increase. The first part of our hypothesis is supported, but the variability decreased, so the second part of the hypothesis was not supported.   There could be several explanations for why there is not much of a declining trend in the "Hard" condition, or why the TFA for epoch 4 is slightly lower than the TFA in epochs 5 and 6: Perhaps three minutes is not enough time for the participants to fully learn the new pattern, or perhaps in epoch 4 participants were able to take a"break" while all of the panes simultaneously reset, and thus had more downtime before responding to the WO openings.</p><p>In any case, our results show that temporal measures of task performance are sensitive to changes in task disruptions and difficulty. There are certainly many more temporal measures to be derived from the raw data output our experimental paradigm can provide, offering a more complete view to the development and changes in human temporal awareness in dynamic tasks, and potentially providing input to humanaware AI agents for their adaptation to human performance.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Limitations</head><p>Due to the homogeneity of our participant population, this experiment should be repeated with a larger, more heterogenous sample. The program output of raw data was also not conducive for derivation of human performance variables. Therefore, development of standard of data output from human-automation interactions for reliable and real-time calculation of valid human performance measures is a critical task for further experiments.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Future Research</head><p>This pilot study showed the feasibility of our experimental paradigm to study human performance in a dynamic task through data unintrusively derived from the task performance. Our experimental program will allow for many additional measures to be recorded, including mouse positions, pane switching before the WO opens (system monitoring), and other timed actions relative to the task. Additionally, there is a good amount of flexibility we have in defining our experimental conditions. Having a longer period per experimental conditions may help us identify trends more clearly and would allow us to have a better view of when someone settles into a pattern. It also could be interesting to incorporate "breaks" into the experiment to measure participants' situation awareness in different ways, for example, by SAGAT <ref type="bibr">(Endsley, 1988)</ref>, or to incorporate SPAM <ref type="bibr">(Durso &amp; Dattel, 2004</ref>) into the task.</p><p>The next steps in this research program will develop an array of measures to be derived from both humans (behavioral, psychophysiological, and subjective) and the system (temporal and cognitive demands of the tasks), and use of machine learning for data fusion (system and human measures) and in classification of human performance into the COCOM modes <ref type="bibr">(scrambled, opportunistic, tactical, and strategic)</ref>. We will also research the transparency of the ML algorithms and robustness of the COCOM classifications across diverse user populations and task demands for their suitability for human-aware AI.</p></div><note xmlns="http://www.tei-c.org/ns/1.0" place="foot" xml:id="foot_0"><p>Proceedings of the Human Factors and Ergonomics Society Annual Meeting 67(1)   </p></note>
		</body>
		</text>
</TEI>
