<?xml-model href='http://www.tei-c.org/release/xml/tei/custom/schema/relaxng/tei_all.rng' schematypens='http://relaxng.org/ns/structure/1.0'?><TEI xmlns="http://www.tei-c.org/ns/1.0">
	<teiHeader>
		<fileDesc>
			<titleStmt><title level='a'>Forecasting Health and Wellbeing for Shift Workers Using Job-Role Based Deep Neural Network</title></titleStmt>
			<publicationStmt>
				<publisher></publisher>
				<date>02/21/2021</date>
			</publicationStmt>
			<sourceDesc>
				<bibl> 
					<idno type="par_id">10289954</idno>
					<idno type="doi"></idno>
					<title level='j'>International Conference on Wireless Mobile Communication and Healthcare MobiHealth 2020: Wireless Mobile Communication and Healthcare</title>
<idno></idno>
<biblScope unit="volume"></biblScope>
<biblScope unit="issue"></biblScope>					

					<author>Han Yu</author><author>Asami Itoh</author><author>Ryota Sakamoto</author><author>Motomu Shimaoka</author><author>Akane. Sano</author>
				</bibl>
			</sourceDesc>
		</fileDesc>
		<profileDesc>
			<abstract><ab><![CDATA[Shift workers who are essential contributors to our society, face high risks of poor health and wellbeing. To help with their problems, we collected and analyzed physiological and behavioral wearable sensor data from shift working nurses and doctors, as well as their behavioral questionnaire data and their self-reported daily health and wellbeing labels, including alertness, happiness, energy, health, and stress. We found the similarities and differences between the responses of nurses and doctors. According to the differences in self-reported health and wellbeing labels between nurses and doctors, and the correlations among their labels, we proposed a job-role based multitask and multilabel deep learning model, where we modeled physiological and behavioral data for nurses and doctors simultaneously to predict participants’ next day’s multidimensional self-reported health and wellbeing status. Our model showed significantly better performances than baseline models and previous state-of-the-art models in the evaluations of binary/3-class classification and regression prediction tasks. We also found features related to heart rate, sleep, and work shift contributed to shift workers’ health and wellbeing.]]></ab></abstract>
		</profileDesc>
	</teiHeader>
	<text><body xmlns="http://www.tei-c.org/ns/1.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink">
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1">Introduction</head><p>Around 20% of the workforce in the world involves in shift work <ref type="bibr">[48]</ref>. Their irregular shift work brings a high risk of poor health and wellbeing. For example, shift work disrupts workers' circadian rhythms and causes problems such as sleep disorder and insomnia <ref type="bibr">[10]</ref>. In addition to the sleep issues, decreased alertness levels were found in healthy shift workers <ref type="bibr">[9]</ref>, which could lead to occupational errors and accidents. Previous studies also showed the potential associations between shift work and pathological disorders such as fatigue, gastrointestinal malfunction <ref type="bibr">[19]</ref>, and an increased risk of colorectal cancer in night shift nurses <ref type="bibr">[38]</ref>. Moreover, more adverse mental health outcomes, emotional exhaustion, and burnout were observed in shift workers compared to daytime workers <ref type="bibr">[5,</ref><ref type="bibr">15,</ref><ref type="bibr">40,</ref><ref type="bibr">43,</ref><ref type="bibr">46]</ref>. In health care domain, physician burnout is estimated to cost 4.6 billion USD per year <ref type="bibr">[13]</ref>.</p><p>To support shift worker's health and wellbeing, monitoring and predicting their day-to-day health and wellbeing trajectories and providing aids to help them prepare for challenging situations might be useful. Besides, mobile devices, such as smartphones and wearable sensors, have become parts of people's daily life, and have been used to detect and predict self-reported health and wellbeing with the help of machine learning models <ref type="bibr">[2,</ref><ref type="bibr">16,</ref><ref type="bibr">21,</ref><ref type="bibr">41,</ref><ref type="bibr">42,</ref><ref type="bibr">49]</ref>. These previous works targeted health and wellbeing detection or prediction as binary classification <ref type="bibr">[2,</ref><ref type="bibr">41]</ref>, 3-class classification <ref type="bibr">[25,</ref><ref type="bibr">49]</ref>, and regression tasks <ref type="bibr">[1,</ref><ref type="bibr">16,</ref><ref type="bibr">49]</ref>. Some of these works developed personalized models by taking participants' demographic information into account <ref type="bibr">[41]</ref> or fine-tuning general models to specific users <ref type="bibr">[50]</ref>. Correlations among self-reported multi-dimensional labels -including subjective mood, health, and stress-were also used in building multilabel neural network models <ref type="bibr">[41]</ref>. In addition, there are some prior works in monitoring shift workers using wearable sensors. Feng et al. extracted a behavioral consistency feature from shift worker wearable data and estimated anxiety levels with an accuracy of 57.8% in binary classification. <ref type="bibr">Mulhall et al. used</ref> sensors integrated in the vehicles to monitor shift workers' eye blinking as a marker of alertness <ref type="bibr">[27]</ref> while driving. Actigraphy has been also used widely for studying sleep for shift work nurses <ref type="bibr">[11,</ref><ref type="bibr">17]</ref>.</p><p>Although these previous works have achieved promising results, there is no work to thoroughly monitor and analyze different job types of shift workers' multidimensional wellbeing and forecast them using machine learning. Furthermore, the models developed previously considered the heterogeneity among participants and correlation among wellbeing labels separately; however, since these two characteristics ubiquitously co-exist, modeling them simultaneously for different job types of shift workers might improve prediction model performance.</p><p>In this work, we collected physiological and behavioral data from hospital shift workers, then we developed machine learning models to predict their next day's wellbeing in binary/3-class classifications and regression tasks. We also verified the rationale of leveraging job role information and multi wellbeing labels simultaneously in the models by analyzing the data. Then, we proposed a multitask multilabel deep learning model that leveraged job role information and correlations among self-reported health and wellbeing labels.</p><p>Our contributions can be summarized as: (i) we collected physiological and behavioral data from hospital shift workers, including nurses and doctors, (ii) we analyzed their physiological and behavioral patterns and found similarities and differences, (iii) we developed a multitask multilabel deep learning model to predict participants' near future wellbeing using wearable sensor, surveys, their job role information, and correlations among wellbeing labels. The details of our proposed model structure, implementation and hyper-parameter information are shared on: <ref type="url">https://github.com/comp-well-org/multitask-multilabelwellbeing-prediction</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2">Related Work</head><p>There are numerous studies on shift workers' health and wellbeing. Heath et al. collected survey data from shift work nurses, and applied statistical analysis in exploring the association among their work shift types, sleep, mood, and diet <ref type="bibr">[14]</ref>. They showed that shift work was significantly negatively related to shit workers' diet, sleep efficiency, and stress levels. Similarly, Books et al. analyzed questionnaire data from shift-working nurses and showed an increased risk of sleep deprivation, family stressors, and mood changes due to the night work shift <ref type="bibr">[3]</ref>.</p><p>In addition, with the rapid development of mobile devices and mobile applications, objective data from wearables and smartphones have been used for studying shift workers. For example, Pereira et al. collected wearable accelerometer data from hospital shift workers and detected 4 levels of their physical activity intensity with an 83% accuracy score <ref type="bibr">[31]</ref>. Feng et al. used wearable devices to collect physiological data from shift work nurses for ten weeks and applied a clustering method for extracting behavioral consistency, which intuitively captures unique behavioral patterns between different groups of nurses <ref type="bibr">[7]</ref>. They further found that behavioral consistency can help predict self-reported work behaviors and anxiety levels. In another work, Feng et al. analyzed physiological and indoor location data from nurses with Fitbit wrist-wearable devices and Bluetooth hubs <ref type="bibr">[6]</ref>. They extracted mutual information features and demonstrated the dependency between an individual's movement patterns and physiological responses.</p><p>Machine learning models have been designed for detecting or predicting health and wellbeing using mobile and sensor data. For example, Bogomolov et al. developed daily stress detection algorithms based on five-month-long weather, mobile phone data (e.g., calls, SMS, and screen usage), and personality survey data from 117 participants <ref type="bibr">[2]</ref>. They obtained stress detection accuracy up to 72% in binary classification tasks. In Moodscope paper, mood (1: negative to 5: positive) was detected with the best mean squared error of 0.229 using the data from the mobile phone and a personalized linear regression model <ref type="bibr">[21]</ref>. Similarly, Asselbergs et al. detected the current mood using mobile phone data with a mean squared error of 0.15 out of -2 to 2 mood scale <ref type="bibr">[1]</ref>. For further improving the model performance, Taylor et al. developed a multitask machine learning model to predict high/low self-reported stress, mood, and health and separately used (i) the demographic information such as gender and personalities of participants and (ii) correlations among labels <ref type="bibr">[41]</ref>. This work also inspired us to use the combination of job role information and label correlations. In this work, we study the differences in daily self-reported health and wellbeing, physiology, and behavior between nurses and doctors, and focus on estimating shift workers' health and wellbeing using the data from mobile sensors and surveys and job-role based deep learning models.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3">Methods</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.1">Data Collection</head><p>Two hundred and forty-one days of multi-modal data were collected from 14 shift workers, including 10 nurses (one male) and 4 doctors (all males) in a hospital in Japan. The average age of all participants was 31.4 years old, with a standard deviation (SD) of 4.2. For each study day, participants wore a Fitbit wristwatch (Fitbit Charge 3) for monitoring their physiological and behavioral activities such as heart rate, sleep, and step counts. The data sampled every 1 min was downloaded from the Fitbit server for data analysis and modeling. In addition, participants filled out daily morning and evening questionnaires to record their behavioral activities, including sleep, work schedule, and caffeinated drinks, alcohol &amp; drug intake.</p><p>Self-reported health and wellbeing labels -including alertness, happiness, energy, health, and stress -were also collected in the morning questionnaire using 0 to 100 scales, with 0 to the most negative and 100 being the most positive (sleepy-alert, sad-happy, sluggish-energetic, sick-healthy, stressed-calm).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.2">Features</head><p>We calculated the following features from the Fitbit data and daily questionnaires:</p><p>Heart Rate. Heart rate and heart rate variability are related to work stress <ref type="bibr">[44]</ref> and mood <ref type="bibr">[39]</ref>. Based on heart rate collected from Fitbit sensor every 1 min, we computed features including daily mean, standard deviation (SD), and entropy of heart rates. We computed sample entropy of heart rate, which represented the self-similarity of a sequence and has been used in physiological time-series data analysis <ref type="bibr">[34]</ref>. To calculate the sample entropy, we first need to set an embedding dimension m. Using the given m, our sequence X with length N can be divided into Nm + 1 sliding windows {X 1 ,...,X N -m+1 }, where X i = {x i , x i+1 , ..., x i+m-1 }. The equation of sample entropy is:</p><p>where</p><p>In our case, the distance d is:</p><p>Generally, m = 2 and r = [0.2 * (SD of X)].</p><p>Sleep. From Fitbit sensors, we obtained sleep duration and sleep efficiency. Then, we calculated the mean and SD values of sleep duration and sleep efficiency across the previous 7, 5, and 3 days. Moreover, using sleep data in one-minute resolution, we calculated sleep regularity with sliding windows across 7 days of participants' data. Sleep regularity is a value of 0-1 based on the likelihood of sleep/wake state being the same time-points 24 h apart, and is associated with health, wellbeing, and academic performance in college students <ref type="bibr">[8,</ref><ref type="bibr">32,</ref><ref type="bibr">37]</ref>. From daily surveys, we obtained a daily feature of the time taken to fall asleep in minutes. Participants also reported how they woke up in the morning: waking up naturally, being awakened by the alarm, or other than alarm. Naps have been shown a positive impact on shift workers' performance, alertness <ref type="bibr">[33]</ref>, and wellbeing <ref type="bibr">[20]</ref>. From participants' questionnaires, we summarized the times and total duration of naps across a day.</p><p>Steps. Total daily number of steps and minute by minute number of steps were recorded in the Fitbit dataset. To measure the variability of participants' physical activities, we computed the mean and SD to indicate step variability across the previous 7, 5, and 3 days. Excluding the sleep time, we counted the minutes of: (i) duration of segments without steps (stationary segments) and (ii) duration of segments with continuous steps (active segments) in 1-min bins. We used the following information entropy equation to calculate the entropy of the two types of physical activity based stationary and active segments:</p><p>where p i represents the probability that the i th item was observed.</p><p>Work. Work schedules and work hours are directly related to symptoms such as sleep disorders and chronic fatigue <ref type="bibr">[4]</ref>. Also, excessive work hours are harmful to workers' health and wellbeing <ref type="bibr">[12]</ref>. We engineered work related features such as daily work shifts, total work duration per day, and overwork duration in minutes according to participants' answers in the questionnaires. There were three different work shifts, and each shift was for eight hours (1: 8:30-16:30, 2: 16:30-0:30, 3:0:30-8:30). Total work duration was actual work time, and the overwork duration was the difference between the actual work hours and the scheduled hours.</p><p>Caffeine, Alcohol and Drug Use. Considering caffeine, alcohol, and drug intake affects workers' alertness <ref type="bibr">[29,</ref><ref type="bibr">35]</ref>, we computed features related to the intake of caffeinated drinks, drug, and alcohol based on the participants' reports: the number of caffeinated drinks per day, and a binary feature for indicating whether the participant had drug or alcohol each day.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.3">Statistical Analysis of Physiological and Behavioral Features Between Nurses and Doctors</head><p>We applied statistical tests to analyze the differences of physiological and behavioral features between two groups, nurses and doctors. Seventy-seven days of data were in the group of doctors, and 164 days of data were in the nurses' group.</p><p>Between 2 groups, we compared the numeric features such as daily average heart rate, steps, and overwork time using Mann-Whitney U test (non-normally distributed features) <ref type="bibr">[23]</ref> and Welch's t-test (normally distributed features) <ref type="bibr">[45]</ref>, whereas the categorical features such as awakening types and working shifts were compared with chi-square test <ref type="bibr">[30]</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.4">Job-Role Based Multitask Multilabel Neural Network</head><p>Neural networks have been widely used in various areas, including face detection <ref type="bibr">[36]</ref>, mood, health, and stress prediction <ref type="bibr">[41]</ref>. These previous outstanding works showed that the design of neural network structure needs the consideration of unique characteristics of data sets used in different applications. As discussed briefly in Sect. 1, in this work, we considered two important aspects: (1) different distributions in health and wellbeing labels based on our participants' demographic information and (2) correlations among health and wellbeing labels. We observed differences in the distributions of self-reported health and wellbeing labels from two job roles, nurses and doctors. Also, there are correlations among the five labels. The details of the data statistics will be discussed in Sect. 5.1.</p><p>To learn different representations corresponding to participant job roles, we applied a multitask learning method, which divided tasks according to participants' job roles. Furthermore, as another form of multitask learning, we used multilabel learning for considering different health and wellbeing labels as tasks. In this way, the model would also fit the correlation among labels. In this work, we designed a job-role based multitask and multilabel neural network model that leveraged user demographic information and correlations among labels at the same time. Figure <ref type="figure">1</ref> shows a simplified version of our model. When training the model, there might be redundant features in our input data that would not help health and wellbeing prediction. In contrast, some non-linear combinations of features might improve our model performance. Thus, we applied a onedimension convolutional neural network (CNN) layer to extract auto-features from our inputs. As shown in Fig. <ref type="figure">1</ref>, we designed convolutional kernels to learn higher-level features across every day feature vectors: 32 row-wise convolutional kernels embedded 32 channels of new features. Then, the CNN extracted features were fed into the multitask neural network. The shared layers in the network learn the representation from all participant data, and the divided branches of the network structure learn the representation independently from participants in different job roles, nurses and doctors. When doctors' data are fed into the model for training, the weights of loss and optimizer of the nurse branch will be set to 0, and vice verse. Furthermore, each branch of the network outputs all five labels (alertness, happiness, energy, health, and stress) from the shared network layers. Therefore, the outputs of our model simultaneously provide the prediction of all five labels for nurses and doctors. The batch loss function of our model can be represented as:</p><p>L ml = l={alert,happy,energy,health,stress} loss(x, y l ) (</p><p>Where x and y represent the input data and the expected output target, respectively. loss is mean squared error loss in regression tasks and cross-entropy loss in the classification tasks. Convolutional neural network kernels are applied for extracting high-level features. Our health and wellbeing prediction is designed for nurses and doctors using a portion of the network trained only using data from either nurses or doctors. Shared layers learn representation from all participants. The final output layers provide the prediction of all five labels simultaneously.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4">Experiments</head><p>Our tasks are formulated in two ways for evaluation: regression and classification tasks. The regression task is to predict the next day's health and wellbeing scores, each in the range of 0-100, whereas the classification task is to predict next day's high/low (binary classification, defined as 100-51, 50-0) or high/mid/low health and wellbeing levels, and high/ mid/low (3-class classification, defined as 100-67, 66-34, or 33-0). Our models use the wearable and survey data up to and including the current day for predicting nurses' and doctors' next day health and wellbeing labels. We compared our job-role based multitask multilabel model (MTML-NN) with following approaches to evaluate the benefits of using demographic information and the correlation among labels: (1) random forest (RF), (2) RBF kernel based support vector machine (SVM), (3) multitask neural network (MT-NN) that used clusters of participants and achieved state-of-the-art performance in a previous study <ref type="bibr">[41]</ref>, (4) multitask neural network with labels as tasks (ML-NN). In addition to applying ML-NN to all participants (ML-NN (all)), we also calculated the prediction results for nurses (ML-NN(N)) and doctors (ML-NN(D)) separately.</p><p>For training and testing our models, we randomly split the dataset into training and testing data in a ratio of 80% to 20%. We applied 10-fold cross-validation and grid search to finalize the hyperparameters for all models mentioned above in the training set. Then, we tested models in the testing set. To make the evaluation process more robust, we repeated the random data split strategy (training/testing : 80%/20%) 10 times to evaluate the model performance. As the evaluating metrics, we use mean absolute errors for the regression models and f1-scores for classification tasks. Furthermore, we adopt focal loss <ref type="bibr">[22]</ref> as the objective function in the classification tasks to mitigate the unbalanced sample size in both binary and 3-class tasks. The Adam optimizer <ref type="bibr">[18]</ref> was used in training the neural networks, with a learning rate of 0.005 and 0.9, 0.999 for &#946; 1 and &#946; 2 .</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.1">Model Weights Analysis</head><p>In addition to the prediction performance, interpretability is also an essential part of machine learning models. Ideally, we would like to provide our prediction results along with reasonable explanations to our participants or health/medical stakeholders. First, from the weights in the RF model, we analyzed the importance of input features. Then, in our deep learning MTML-NN model, we analyzed the importance of the features by examining the parameters in the first CNN layer before the non-linear activation function. Since the CNN kernel we designed is in one-dimension with a size of the number of features, and parameters in the CNN kernel would correspond to the input features. We calculated the average value of each feature on all channels to check the importance of the features. Also, we computed the correlations between the output of the CNN layer and the input features. Features that have higher correlations with the CNN outputs would also be considered important features.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5">Results and Discussion</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.1">Data Statistics</head><p>As shown in Table <ref type="table">1</ref>, the average score of alertness label was the lowest among all five labels; while the stress label (0: pressure-1: calm) showed the highest average score. Compared with other labels, the SD of happiness score was lower. Moreover, the distribution of health and wellbeing labels for nurses and doctors were different. For example, doctors generally had higher subjective alertness and energy than nurses in the morning. In addition, we computed correlations among the five health and wellbeing labels. Figure <ref type="figure">2</ref> shows the correlation coefficients matrix of all labels, and there are different degrees of correlation among the labels. The Pearson test <ref type="bibr">[26]</ref> showed that all five labels were significantly correlated. The linear fitting coefficient of determination (r 2 ) values <ref type="bibr">[28]</ref> between the alert label and other labels ranged from 0.19 to 0.28, while the r 2 values among the happy, energy, health and stress labels were all higher than 0.55, with the highest value being 0.70 (happy and stress).</p><p>We also compared feature distributions between nurses and doctors (Table <ref type="table">2</ref>). We found that the mean heart rate of doctors was significantly higher than that of nurses; whereas the variability of heart rate, defined as SD and sample entropy, was higher in nurses than doctors. In terms of sleep, we found that doctors showed higher sleep efficiency and lower sleep irregularity than nurses. Further, We found statistical differences between nurses and doctors in movement features, including mean and SD of daily steps across the previous 7 days, and the entropy for stationary/active segments. We did not observe any statistical differences between nurses and doctors in working shifts and total work hours among shift work features. However, we found that overwork was more common among doctors. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.2">Wellbeing Prediction</head><p>The classification and regression performance using different models is shown in Table <ref type="table">3</ref>. Our proposed job-role based MTML-NN performed the best for four labels in binary classification and all wellbeing labels in 3-class classification and regression (ANOVA, Tukey, p &lt; 0.05). Our results showed the benefits of our proposed simultaneous job role and correlated label modeling, especially in  3-class classification and regression. However, according to the performance in 3class classification, we found poor classification performance for some classes. For example, in the 3-class alertness classification, the high-alertness class precision and recall values were only 0.16 and 0.27 in respectively; and our low-energy class prediction was also relatively low with a precision of 0.33 and a recall of 0.24. These errors might come from the data imbalance problem. In the 3-class classification tasks, the high alertness labels accounted for only 20% of all labels, and the low energy labels accounted for 15% of all labels. Furthermore, we also found the benefits of using job role information or multiple labels separately. For example, in the alertness prediction, job-role based MT-NN showed significant improvement from NN for both binary and 3-class classification. Besides the overall f1-score, we observed some improvements revealed in each class. For example, in the 3-class alertness classification tasks, MT-NN model provided significantly higher recall and precision scores in low and middle alertness classification compared to NN model (Welch's t-test, p &lt; 0.05). We did not observe any significant improvement in the regression tasks. However, the average prediction MAE of MT-NN was lower than that of NN. Significant improvements were observed in ML-NN compared to NN in almost all tasks. For example, in the regression tasks, the ML-NN (all) performed statistically significantly better than NN in predicting alertness, happiness, energy, and stress labels.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="5.3">Weight Analysis</head><p>From the RF model, for both the binary and 3-class classification happiness prediction tasks, we found features including mean heart rate and heart rate sample entropy across the day, sleep duration, sleep regularity, and the SD of sleep efficiency across the previous seven days, were the most important. In the alertness prediction tasks, work shifts, stationary segment entropy, mean step, and mean sleep duration across the previous 7, 5 days played important roles. The analysis of the parameters in the CNN layer in the MTML-NN model and the correlations between the CNN output and input features indicated that features including heart rate sample entropy, sleep regularity, sleep efficiency, work shifts, steps, and active segment entropy -contributed to health and wellbeing prediction. For example, from the correlation analysis, we found that the sleep efficiency, sleep regularity, and daytime work shift were positively related to the wellbeing (Pearson test, p-value &lt; 0.05/(# of features)); whereas the step and the entropy of active segments were negatively related to the wellbeing (Pearson test, p-value &lt; 0.05/(# of features)). Our findings were consistent with some prior results. For example, according to the previous works, sleep influences physical and psychological health <ref type="bibr">[47]</ref>, and stress <ref type="bibr">[24]</ref>; sleep regularity is associated with mood <ref type="bibr">[37]</ref>. Previous studies also indicated the association between work shifts and stress levels <ref type="bibr">[46]</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="6">Conclusion</head><p>In this work, we collected physiological and behavioral wearable sensor data as well as survey data from shift-work nurses and doctors, and compared their physiology and behaviors between two job roles. Then, we proposed a job-role based multitask and multilabel learning model structure to predict shift workers' health and wellbeing for next day using sensor and questionnaire data. The proposed model outperformed the benchmark models, including RF and SVM as well as the previous state-of-the-art models. The analysis of model weights showed that health rate, work shifts, sleep parameters such as sleep regularity and sleep efficiency contributed to shift workers' health and wellbeing labels. As future work, we will collect more data from shift workers and design a system to improve shift workers' health and wellbeing.</p></div></body>
		</text>
</TEI>
