<?xml-model href='http://www.tei-c.org/release/xml/tei/custom/schema/relaxng/tei_all.rng' schematypens='http://relaxng.org/ns/structure/1.0'?><TEI xmlns="http://www.tei-c.org/ns/1.0">
	<teiHeader>
		<fileDesc>
			<titleStmt><title level='a'>VR Sickness Versus VR Presence: A Statistical Prediction Model</title></titleStmt>
			<publicationStmt>
				<publisher></publisher>
				<date>01/01/2021</date>
			</publicationStmt>
			<sourceDesc>
				<bibl> 
					<idno type="par_id">10251038</idno>
					<idno type="doi">10.1109/TIP.2020.3036782</idno>
					<title level='j'>IEEE Transactions on Image Processing</title>
<idno>1057-7149</idno>
<biblScope unit="volume">30</biblScope>
<biblScope unit="issue"></biblScope>					

					<author>Woojae Kim</author><author>Sanghoon Lee</author><author>Alan Conrad Bovik</author>
				</bibl>
			</sourceDesc>
		</fileDesc>
		<profileDesc>
			<abstract><ab><![CDATA[Although it is well-known that the negative effects of VR sickness, and the desirable sense of presence are important determinants of a user's immersive VR experience, there remains a lack of definitive research outcomes to enable the creation of methods to predict and/or optimize the trade-offs between them. Most VR sickness assessment (VRSA) and VR presence assessment (VRPA) studies reported to date have utilized simple image patterns as probes, hence their results are difficult to apply to the highly diverse contents encountered in general, real-world VR environments. To help fill this void, we have constructed a large, dedicated VR sickness/presence (VR-SP) database, which contains 100 VR videos with associated human subjective ratings. Using this new resource, we developed a statistical model of spatio-temporal and rotational frame difference maps to predict VR sickness. We also designed an exceptional motion feature, which is expressed as the correlation between an instantaneous change feature and averaged temporal features. By adding additional features (visual activity, content features) to capture the sense of presence, we use the new data resource to explore the relationship between VRSA and VRPA. We also show the aggregate VR-SP model is able to predict VR sickness with an accuracy of 90% and VR presence with an accuracy of 75% using the new VR-SP dataset.Index Terms-VR sickness assessment (VRSA), VR presence assessment (VRPA), natural video statistics (NVS), human visual system (HVS).
I. INTRODUCTIONI N RECENT decades, there have been tremendous advances in the development of virtual reality (VR) technologies [1]. VR devices have been successfully deployed as a way of simulating real-world experiences in a wide variety of domains including gaming, simulators and medical clinics, using increasingly lightweight, comfortable, and immersive head-mounted displays (HMD). However, the quality of experience (QoE) of VR users is often severely reduced by VR sickness. In many ways, providing a satisfactory and realistic sense of presence to VRs users involves navigating trade-offs between immersion and comfort. Manuscript]]></ab></abstract>
		</profileDesc>
	</teiHeader>
	<text><body xmlns="http://www.tei-c.org/ns/1.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink">
<div xmlns="http://www.tei-c.org/ns/1.0"><p>Here we focus on the development of ways to navigate these trade-offs. Specifically, we develop a model of VR sickness assessment (VRSA) and also develop a VR presence assessment (VRPA) model that takes advantage of predicted VR sickness scores, inspired by the trade-off relationship between VR sickness and VR presence. This allows the model to predict an optimal combination of factors contributing to satisfaction in VR.</p><p>In many previous studies, VR sickness has been reported to induce oculomotor symptoms such as visual fatigue and difficulty focusing; physiological reactions including burping, salivation and sweating; and disorientation, dizziness, and vertigo <ref type="bibr">[2]</ref>- <ref type="bibr">[4]</ref>. One of the most common causes of VR sickness is a sensory mismatch between the vestibular system and the visual system. The vestibular system, which perceives actual movement, and the visual system, which perceives projected motion fields, may provide conflicting visual percepts <ref type="bibr">[5]</ref>, <ref type="bibr">[6]</ref>. When there is a visual perception of movement, especially self-motion, that is experienced in a static physical state, visual-vestibular conflicts may arise in the brain. Because of this, a number of authors have reported ways of measuring VR sickness by quantifying the amount of perceptual motion. For example, the authors of <ref type="bibr">[7]</ref> predict VR sickness by calculating differences between perceived motions and head movements using a visual-vestibular conflict model. In <ref type="bibr">[8]</ref>, a deep autoencoder based model is developed that predicts exceptional motion. The model is also generalized using a generative adversarial network (GAN) <ref type="bibr">[9]</ref>.</p><p>However, these studies have not benefited from the availability of sizeable labeled datasets. We utilize a new, large database that we have created to develop a new, highly competitive model that is capable of predicting both the degree of VR sickness that may be felt, as well as the VR sense of presence. This dual model may prove useful for mediating these negative and positive aspects during content creation or display.</p><p>Our approach to VRSA differs from prior methods in that it seeks to quantify losses of statistical regularity of VR content that are predictive of sickness arising from VR content. Specifically, the contributions that we make are summarized:</p><p>1) We built the first large database addressing both VR sickness and VR presence, including subjective labels. 2) We develop new predictive models for VRSA and VRPA that are based on natural VR content statistics. 3) We demonstrate state-of-the-art performance of the new models for VR sickness and VR presence prediction.</p><p>The remainder of the paper is organized as follows. Section II introduces related works on VR sickness and presence prediction. Section III describes the design of the proposed VR-SP database. Section IV introduces the VRSA based on predictive statistical features, while Section VI describes our VRPA model, which uses the predicted VR sickness scores, along with measures of visual activity and content. Section V discusses experimental results, and conclusions are given in Section VII.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>II. RELATED WORKS</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>A. Sickness Assessment in VR Environment</head><p>There have been several studies of VR sickness from a physiological or behavioral perspective <ref type="bibr">[10]</ref>- <ref type="bibr">[13]</ref>. Most of these have involved simulator sickness questionnaires (SSQ) or physiological tests such as electroencephalogram (EEG), galvanic skin response (GSR), electrogastrogram (EGG), or heart rate [10]- <ref type="bibr">[12]</ref>. However, these sensor modalities are quite noisy, and it is difficult to engage subjects in viewing and rating a large amount of VR contents under these protocols. Moreover, these tests have typically deployed simple visual patterns (e.g., stripes) <ref type="bibr">[6]</ref>, <ref type="bibr">[14]</ref>, <ref type="bibr">[15]</ref>.</p><p>However, content that is being viewed on modern VR devices, such as the Oculus Rift/Quest and HTC-VIVE HMDs, is trending towards increased realism and naturalness, along with more immersive experiences. Being able to automatically and objectively predict VR sickness in HMD environments has become highly desirable.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>B. Natural Scene Statistics</head><p>A number of natural scene statistics (NSS) methods have been developed to measure losses of statistical regularities arising from distortions of picture and videos <ref type="bibr">[16]</ref>. The efficacy of this approach has been amply demonstrated by the development of highly distortion-sensitive bandpass decomposition methods in the wavelet domain <ref type="bibr">[17]</ref>, the discrete cosine transform (DCT) <ref type="bibr">[18]</ref>, and in spatial <ref type="bibr">[19]</ref> and frame difference <ref type="bibr">[20]</ref> domains. Inspired by these advances, we have developed methods for predicting VR sickness from variations in immersive video statistics.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>C. Sense of Presence</head><p>There has been a variety of studies of the sense of presence experienced when viewing 2D and stereoscopic 3D (S3D) visual content. Early on, Ijsselsteijn et al., introduced the concept of visual presence and how to measure it <ref type="bibr">[21]</ref>. They also conducted a study of subjective presence on S3D content, and derived objective measures of the sense of presence <ref type="bibr">[22]</ref>. They also gathered subjective opinions of both presence and discomfort of viewers of stereoscopic cinema. Oh and Lee <ref type="bibr">[23]</ref> developed a visual presence assessment model utilizing the geometry of S3D displays.</p><p>In connection with VR sickness, it is generally thought that VR sickness is one of the most crucial factors causing users to feel differences between virtuality and reality <ref type="bibr">[24]</ref>- <ref type="bibr">[28]</ref>. When VR sickness in virtual space exceeds that experienced in the real world, then the sense of realism is reduced <ref type="bibr">[29]</ref>. However, the relationship between VR sickness and VR presence has not yet been quantified adequately, because of the lack of a sufficiently large VR content database with subject labels of presence and sickness. Because of this, the problem of predicting VR presence and VR sickness, and the relationship between them, remains unresolved.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>III. VR-SP DATABASE</head><p>In this section, we describe the new VR-SP database. It is well known that both VR sickness and presence are regarded as critical factors that affect the popularity and viability of VR products. Existing databases of psychometric scores are quite limited in the number of included VR contents <ref type="bibr">[8]</ref>, <ref type="bibr">[9]</ref>, <ref type="bibr">[30]</ref>. Moreover, most of these databases focus only on VR sickness, with little or no analysis of VR presence or its relationship with VR sickness. Towards advancing progress in this direction, our new VR-SP database contains a wide variety of diverse VR source contents, along with corresponding subjective scores on VR sickness and VR presence. The VR-specific content in the database can be viewed on any popular VR device, such as the HTC-VIVE or the Oculus Rift.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>A. VR Content</head><p>The new VR content database contains 10 reference VR scenes providing a variety of virtual experiences (e.g., VR games, roller-coaster, and outer space), as tabulated in Table <ref type="table">I</ref>. To create diverse experiences of VR sickness, highly diverse camera rotations (yaw, pitch, and roll) and translations (forward, backward, and lateral) were deployed, depending on each scene. All the reference contents were implemented with various presets on the Unity platform.</p><p>Since our goal is to design a database enabling the analysis and quantification of VR sickness and VR presence, a total of 100 VR contents were created from the 10 reference VR scenes. Toward this, each reference content was diversified to 10 variations by combining two types of motion, and four levels of velocity, applied to the 10 reference VR contents in Fig. <ref type="figure">1</ref>. In addition to spatially diversifying the VR contents, we reduced the level of detail in the scenes. The detail-reduced scenes were those containing only the two lower levels of velocity. Fig. <ref type="figure">2</ref> diagrams the variations (V1-V10) of each reference scene. As shown in the figure, the motion types and velocity levels were applied to the individual reference and detail-reduced contents. The reference videos have detailed environments, as shown in Fig. <ref type="figure">3</ref>, while the modified scenes contain less detail. The VR-SP database also contains two types of motion. The first motion category is a simple movement, such as linear motion, while the second includes complex motions, such as rotations and rapid transitions. Each motion type includes four levels of velocity, while for the detail-modified scenes, only motion type 1 was used at velocity levels 1 and 2.</p><p>After processing them with the various motion types and velocity levels mentioned above, a subjective study was conducted, wherein all of the VR sequences were viewed    </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>B. Subjective Assessment</head><p>As recommended in the ITU-R BT.500-13 <ref type="bibr">[31]</ref> and BT.2021 <ref type="bibr">[32]</ref> standards, we adopt a single-stimulus subjective scoring methodology for both sickness and presence assessments. The 21 inexperienced subjects (satisfying the subject criteria recommended in <ref type="bibr">[31]</ref>) were of ages ranging from 22 to 34 years. All of the subjects were screened for normal visual acuity on the Landolt chart. The overall protocol included four evaluation sessions, each containing 25 randomly shuffled VR contents from the database. After each individual session, rest periods of 5 min duration were inserted, to minimize any accumulated feeling of VR sickness <ref type="bibr">[33]</ref>. In each session, a VR sequence was displayed for 9 seconds, after which the subject assessed their feelings of sickness and presence. Similar to previous work <ref type="bibr">[34]</ref>, VR presence is defined here as the subjective sense of presence ("actually being there") in the user's space or environment. To measure VR sickness, a simple scale of experienced sickness was used rather than using a lengthy questionnaire, which would have subjected the subjects to unnecessarily long sessions (several hours), and would have made it impossible to collect large amounts of data. This was facilitated by a user controller-based VR interface that allowed the subjects to interactively score in the same display environment. The subjects rated the contents on a discrete, 5-points Likert scale marked as follows.</p><p>We and others <ref type="bibr">[35]</ref> have found that viewers do experience VR sickness on such short time scales, and indeed early sensations of sickness may be the most important to detect, before they accumulate and become more severe.</p><p>The VR Sickness labels on the scale were: Extremely Uncomfortable (5), Uncomfortable (4), Mildly Comfortable (3), Comfortable (2), and Very Comfortable (1). The VR Presence labels were Excellent (5), Good (4), Fair (3), Poor (2), and Bad <ref type="bibr">(1)</ref>.</p><p>The parenthetical numbers indicate the numerical associations with the labels. After all the subjects scored the videos, the mean opinion scores (MOS) of both sickness/presence were obtained. Fig. <ref type="figure">4</ref> plots the distribution of MOS values of both VR sickness and VR presence. As shown in the figure, the distribution is not biased to a specific score for each target value. We summarize major information about the test environment in Table <ref type="table">II</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>C. Trade-Offs Between VR Sickness and VR Presence</head><p>Fig. <ref type="figure">5</ref> shows a scatter diagram and fitted curve of the MOS scores of VR sickness and VR presence plotted against each other, over all samples on the VR-SP database. Interestingly, the VR presence score is gradually lowered in the very low VR sickness region (MOS &lt; 0.4). Furthermore, in the region with excessively high VR sickness (MOS &gt; 0.8), the VR presence scores also trended gradually lower. However, most of the areas with high VR presence were accompanied by moderate VR sickness scores. This relationship may be more clearly seen by the fitted curve. By observation, the most intuitive way to predict VR presence is to discover a specific range where the user experiences moderate VR sickness and proper immersion. Therefore, we will use the predicted VR sickness score as a feature of VRPA in Section V.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>IV. PROPOSED VR SICKNESS ASSESSOR</head><p>Most VRSA studies have focused on understanding VR sickness, by using motion estimation to probe the level of visual-vestibular sensory conflict. Here we instead use natural video statistic (NVS) models derived in the context of a set  of perceptually relevant processes. Moreover, we devise an exceptional motion feature using a measurement of the correlation between statistical features drawn from instantaneous and averaged temporal processing. The design of our VR sickness model is inspired by prior NVS approaches <ref type="bibr">[20]</ref>, <ref type="bibr">[36]</ref> whereby the losses of statistical regularities arising from undesirable visual content, such as distortions. The overall flow of the VRSA framework is depicted in Fig. <ref type="figure">6</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>A. Preprocessing</head><p>To realize the most vivid VR experiences, providing a wide field of view (FoV) is an important factor. Most HMD devices are designed with rectangular displays via a lens system, hence barrel distortion is unavoidable, leading to geometric distortions along the image border.</p><p>For example, Fig <ref type="figure">7</ref> shows an example of a rendered VR content, and its distortion corrected version. As shown in Fig. <ref type="figure">7</ref> (a) (red boxes), borderline pixels are stretched compared to those in the central regions. Accordingly, when a motion estimation method or an optical flow algorithm is applied, prediction errors arising from lens distortion occur. To overcome this, we geometrically correct each rendered frame via a pincushion transformation, as shown in Fig. <ref type="figure">7</ref> (b) (See red boxes in Fig. <ref type="figure">7 (b</ref>)) <ref type="bibr">[37]</ref>. This also matches the corrected image that viewer sees. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>B. Spatio-Temporal Frame Differences</head><p>After preprocessing the VR sequence, we have luminance frames {F 1 , F 2 , . . . , F T } of dimensions N H&#215;W&#215;T , where H, W and T are the height, width and total number of frames, respectively. Given each VR video, we do not explicitly estimate motion, but instead compute temporal changes of luminance. As discussed in <ref type="bibr">[38]</ref> temporal frame differences obey regular statistical laws, while optical flow/motion vectors are much less regular. The time-averaged p-norm frame difference is defined as:</p><p>where x represents a tuple of spatial indices (x = {x 1 , x 2 },</p><p>, over a set of consecutive frames indexed t.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>C. Rotational Frame Differences</head><p>From previous studies, it is well known that rotational image motions, such as roll or pitch, are highly related to VR sickness <ref type="bibr">[39]</ref>, <ref type="bibr">[40]</ref>. One feasible strategy to model rotational statistics is to reconstruct rotational representations using the frame difference maps, then estimate statistical features on them. Fig. <ref type="figure">8</ref> shows nine directional modes with corresponding angles and their reconstructed rotational motion maps (i.e., roll motion, zoom-in/out, horizontal and vertical motions). Fig. <ref type="figure">8</ref> (a), we first calculate spatially displaced frame differences over frames, depending on the directional mode. As shown in <ref type="bibr">[41]</ref>, spatially-displaced frame differences are an effective way to capture space-time statistical regularities predictive of distortion. At each frame index t and spatial index x, spatially displaced frame difference maps are computed as</p><p>where w(&#8226;; a &#952; ) is a warping function over the vector of mode parameters dictated by the angle &#952; . We use eight mode directions to represent motion elements, thus  ). Let F denote the spatio-temporal frame difference map, obtained by the same procedure in <ref type="bibr">(1)</ref>. Then the eight spatially displaced frame difference maps with their corresponding modes are</p><p>To represent a rotational frame difference, first divide the difference map of each mode into nine local patches. Then, as shown in Figs. <ref type="figure">8 (b</ref>)-(e), the corresponding local patch of each mode is reconstructed to capture roll, zoomin/out, horizontal and vertical motions (Figs. <ref type="figure">8 (b)&#8594;</ref></p><p>). Each rotational motion includes directional coefficient pairs, and the final rotational frame difference is defined as the sum of each pair:</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>D. Natural Video Statistics</head><p>1) Frame Difference Normalization: Given a frame differenced sequence F t , the mean subtracted contrast normalization (MSCN) coefficients of it are obtained as <ref type="bibr">[42]</ref>:</p><p>), and over a set of consecutive frame time samples (frame indices), t &#8712; [1, T ] where</p><p>and</p><p>denote the weighted local mean and weighted contrast of each frame difference map, respectively, where w k,l is a Gaussian weighting function sampled out to 3 standard deviations and rescaled to unit volume over (k = -K , . . . , K ), (l = -L, . . . , L). In our experiments, we fixed the semi-saturation constant in the divisive normalization to C = 0.01 and took K = L = 9. The same MSCN process is also applied to each spatial frame F t , yielding normalized coefficients Ft . Here after we will drop the temporal index t and the spatial indices x in (1)-( <ref type="formula">5</ref>)</p><p>2) Statistical Characterization: To characterize the statistical features from the normalized MSCN maps, we use the parametric asymmetric generalized Gaussian distribution (AGGD) <ref type="bibr">[43]</ref>:</p><p>To estimate the parameters of the AGGD (&#956;, &#947; , &#946; l , &#946; r ), we fit it to the histograms of MSCN coefficients over T frames on each VR video sample (and on frames, displaced frame differences, and rotated frame differences), using the popular moment-matching based approach in <ref type="bibr">[43]</ref>.</p><p>In this way, we extract various statistical features on five frame difference maps: F p , R p , Z p , H p , and V p (all of which are functions of (x,t), with notation dropped for  brevity), where the norm index in 1 is either p = 1 or p = 2. Specifically, after applying MSCN normalization, 11 feature maps are obtained for each frame on the spatio-temporal frame difference map &#710; F p and the normalized rotational frame difference maps ( &#710; R p , &#710; Z p , &#710; H p and &#710; V p ) using two norms ( p &#8712; {1, 2}) in addition to the normalized spatial frame F. Finally, we extract statistical features (&#956;, &#947; , &#946; l , &#946; r ) from each spatial feature map, denoting them collectively as &#958; k , k &#8712; {1, . . . , 4}, and likewise extract spatio-temporal features &#966; l p , l &#8712; {1, . . . , 4} on normalized frame differences, and rotational features &#948; m p , m &#8712; {1, . . . , 16} on normalized, rotated frame differences. Furthermore, Fig. <ref type="figure">10</ref> shows 3D scatter plots of the (logarithms of the) extracted parameters &#947; , &#946; l , &#946; r obtained by the fitting process in <ref type="bibr">(6)</ref>, on the five different frame difference maps ( F, R, Z , H and V ), on all the VR sequence samples in the VR-SP database. Each sample is colored as closer to purple as the feelings of VR sickness MOS increased, and towards yellow with decreased feelings of VR sickness. It may be seen that the samples vary with the degree of sickness, suggesting that the extracted features can play an important role in predicting VR sickness.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>E. Exceptional Motion Feature 1) Instantaneous vs. Averaged Temporal Correlation:</head><p>Prior VR sickness studies <ref type="bibr">[44]</ref>, <ref type="bibr">[45]</ref>, have suggested that the visual-vestibular conflict is strongly affected by exceptional motions such as rapid turning movements, e.g., a user experiencing VR content of driving a racing car on a road with many roadblocks. An intuitive way to model exceptional motion is to compare temporal motion flow to spatially local motion flow. To do this, we utilize instantaneous vs. local temporal correlation feature, whereby we measure the correlations between features drawn from instantaneous normalized coefficients (i.e., spatio-temporal features &#966; l p ) and locally averaged temporal normalized coefficients.</p><p>2) Temporal Normalization: We slightly modify the statistics of the MSCN normalization (3) to characterize averaged temporal statistics over multiple frames. Using the subscript Fig. <ref type="figure">10</ref>. 3D scatter plots of (logarithms of) shape, left scale and right scale obtained by fitting the AGGD model to (a) the spatio-temporal frame differences, (b) the roll motion frame differences, (c) the zoom-in/out motion frame differences, (d) the horizontal motion frame differences and (e) the vertical motion frame differences, using all the VR video samples from the VR-SP database. Purple colors indicate MOS associated with greater feelings of sickness, while yellow samples indicate MOS values associated with lesser feelings of sickness.</p><p>'MF' to denote multi-frame, define</p><p>and</p><p>where F t (x) is as before, over a set of consecutive frame time samples t &#8712; [1, T ]. Then the multi-frame MSCN normalization Ft MF is computed by using &#956; MF and &#963; MF as in <ref type="bibr">(3)</ref>. We compute all the same parameters as before, and denote the overall feature vector as p = {&#968; l p ; l = 1, . . . , 4}.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>3) Correlation Between Instantaneous and Averaged Coefficients:</head><p>The spatio-temporal features &#968; l p , l &#8712; {1, . . . , 4} of FMF are obtained by the same AGGD fitting procedure as in Section IV-D. The correlation feature &#223; p between the instantaneous and averaged feature vectors ( p = {&#966; l p } and p = {&#968; n p }) is</p><p>where p is the norm parameter p &#8712; {1, 2}, and &#956;(&#8226;) and &#963; (&#8226;) denote the average and variance of each feature vector, respectively. Fig. <ref type="figure">11</ref> plots the histograms of the correlation feature &#223; 1 for low sickness MOS (MOS &lt; 0.2) and high sickness MOS (MOS &gt; 0.7) instances on the VR-SP database. As may be seen, the correlations are widely distributed for low sickness VR sequences, as compared to high sickness VR sequences.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>V. APPLICATION: VR PRESENCE ASSESSOR</head><p>As mentioned earlier in Section III (also shown in Fig. <ref type="figure">5</ref>), there are strong non-linear trade-offs between VR sickness and VR presence. Therefore, we use the predicted VR sickness scores produced by VRSA as a feature, to help estimate the degree of VR presence. As shown in Fig. <ref type="figure">5</ref>, the VR presence scores (vertical axis) are distributed with a large variance. Therefore, to accurately estimate VR presence, a number of presence-directed features are used to drive the VRPA model <ref type="bibr">[46]</ref>, <ref type="bibr">[47]</ref>: a visual activity feature, and several content features (luminance gradient, color gradient, luminance saturation, and color saturation).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>A. Additional Features for VR Presence</head><p>Fig. <ref type="figure">12</ref> shows the same scatter diagram as Fig. <ref type="figure">5</ref>, but with each plotted point color-coded by the values of each of the additional features. Specifically, in Fig. <ref type="figure">12 (a)-(e)</ref>, the color of each sample indicates larger feature values close to yellow, and lower feature values are closer to purple. It may be seen that the distribution of each feature is nicely divided in the vertical direction. Each feature is calculated as follows:</p><p>1) Visual Activity: Visual activity has previously been used to quantify visual comfort, preference and presence based on statistical analyses of visual content feature maps <ref type="bibr">[46]</ref>, <ref type="bibr">[47]</ref>. To calculate visual activity, first compute the content feature maps (luminance/color gradient and saturation). Then, the normalized luminance/color gradient map is obtained as and</p><p>where G l and G c are the normalized luminance and color gradients, and u and v are horizontal and vertical indices, respectively. I is the luminance map, and G m l is the maximum luminance gradient of I , which is used as a normalization factor. C a and C b are color maps computed after converting the images into the perceptually uniform CIELab color space <ref type="bibr">[48]</ref>. G m c,a and G m c,b are the maximum color gradients of C a and C b , which are used for normalization. Fig. <ref type="figure">13 (b</ref>) shows an exemplar luminance gradient map. In addition, the normalized luminance/color saturation is given by</p><p>and</p><p>where S ca and S cb are the normalized color saturations on color spaces C a and C b , respectively. Then the normalized color saturation is represented by S c = S ca &#215; S cb . When the luminance/color saturation values are towards brighter or darker, then the saturation values approach 1. Fig. <ref type="figure">13</ref> shows an example of a normalized luminance saturation map. A discrete wavelet transform is performed on each normalized luminance/color gradient and saturation map. Then, the histogram of the wavelet coefficients of each subband are fitted with a generalized Gaussian distribution (GGD) <ref type="bibr">[49]</ref>. Let &#947; k,i denote the GGD shape parameter of the i th wavelet sub-band of the k th feature map (k &#8712; {G l , G c , S l and S c }).</p><p>We use the shape parameter as a major descriptor of visual activity, since it captures the distribution of energy across the wavelet-transformed feature maps <ref type="bibr">[47]</ref>. Using the shape parameter, the visual activity A k,i of the i th sub-band of the k th feature map is defined as:</p><p>where the decomposition level set to 2 the number of sub-bands i is set to 7; c k , c k,D and c k,U are normalization factors on the k th feature map. Since the range of visual activity varies with the feature maps, the normalization factors c k , c k,D and c k,U are applied on the k th feature map, where c k is the reference operating point of the k th feature map, and c k,D and c k,U are the trailing and leading edges of the region over which ( <ref type="formula">14</ref>) is approximately linear <ref type="bibr">[47]</ref>.</p><p>In our experiments, the normalization factors were obtained by in <ref type="bibr">[50]</ref>,</p><p>Finally, we calculate the visual activity of each feature map by aggregating them over all sub-bands:</p><p>where A k is the activity of the k th feature map having N sub sub-bands:</p><p>Here, w k is an empirical weight on each feature map.</p><p>2) Content Features: We summarize the content feature maps (CFs) using three different pooling methods. The first is the mean value of (each) feature map, while the other two features are computed by calculating the upper and lower spercentiles % (mean of upper/lower s% of each feature map,), respectively. This yields 12 pooled CF features.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>B. Overall Feature Configurations</head><p>We use the features defined above in Sections IV-V and the VR sickness/presence MOS scores to train a support vector regression (SVR). The extracted VRSA features can be categorized as spatial features (SF), spatio-temporal features (STF), rotational features (RF) and exceptional motion features (EMF). For VRPA, there are sickness features (SCF), visual activity features (AF), and content features (CF). Overall, there are 46 VR sickness features and 14 VR presence features.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>VI. EXPERIMENTAL RESULTS</head></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>A. Experimental Setup</head><p>To validate the performance of VRSA and VRPA, we employed two standard measures: Pearson's linear correlation coefficient (PLCC) <ref type="bibr">[52]</ref>, and Spearman's rank-order correlation coefficient (SROCC). A value close to 1 for SROCC and PLCC indicates better performance. We followed the validation strategy in <ref type="bibr">[53]</ref>- <ref type="bibr">[57]</ref>: first, we randomly divided the VR video database (reference and training) into two content-separated subsets (80% for training and 20% for testing). An SVR was utilized as the regression tool, since it has demonstrated excellent performance on other high-dimensional regression problems, such as 3D discomfort assessment and visual preference prediction <ref type="bibr">[46]</ref>. To implement the SVR, we used the libSVM package <ref type="bibr">[58]</ref> with the radial basis kernel.</p><p>The correlation results that we report are the median correlation over 2000 iterations of randomly dividing the training and testing sets to eliminate any biases (cross-validation).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>B. Dataset</head><p>We benchmarked the VRSA prediction model on two different datasets: ETRI-VR <ref type="bibr">[30]</ref> and the new VR-SP dataset. The VRPA prediction model was tested on the VR-SP dataset. The ETRI-VR database includes 2 reference VR contents with 52 slightly modified scenarios, along with corresponding subjective sickness scores <ref type="bibr">[30]</ref>. As mentioned earlier, the VR-SP dataset includes 10 reference VR contents and 100 variations of the reference VR contents with their corresponding sickness and presence MOS. All sickness MOS values in these databases are scaled to [0, 1], where 1 represents to the extreme discomfort.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>C. VRSA Prediction Results</head><p>We compared the predictions of VRSA against five VR sickness assessment models: the optical flow-based method in <ref type="bibr">[51]</ref>, VRSP <ref type="bibr">[7]</ref>, Kim et al. <ref type="bibr">[30]</ref>, Kim et al. <ref type="bibr">[8]</ref> and VRSA-NET <ref type="bibr">[9]</ref>. For the optical flow-based method, the model was implemented using an estimated average optical flow magnitude <ref type="bibr">[51]</ref>. For Kim et al. <ref type="bibr">[8]</ref> and VRSA-NET <ref type="bibr">[9]</ref>, the experimental setup was the same as in the original work. In the case of Kim et al. <ref type="bibr">[30]</ref>, since our database does not include brain signals, we only utilized visual features from the combined CNN-RNN network.</p><p>1) Benchmark Results: Table <ref type="table">III</ref> tabulates the performance comparison over the tested models on the VR-SP and ETRI-VR databases. When measured on all databases, the standard deviations of PLCC and SROCC after 2000 trials were less than &#8764;0.03. As shown in Table, VRSA delivered significantly better predictive performance than the other models in terms of both correlation and reliability. We also conducted a t-test of statistical significance on the SROCC values over 50 trials of all pairs of benchmark models. Table <ref type="table">VI</ref> shows the results of the t-tests on the VR-SP database. The symbols "1", "0" and "-1" indicate that the performance of the model in the row is statistically better, indistinguishable, or worse, respectively, than the compared methods in the column. We set the confidence level to be 95% (i.e., significance is determined the p-value is less than 0.05). As shown in the results, VRSA was more predictive of VR sickness than the other benchmark models with statistical significance.</p><p>2) Performance on Individual Motion Types: Since the VR-SP database contains a variety of motion types, we also tested the model according to these types. Table V reports the SROCC of the compared VRSA algorithms on the VR-SP database against motion type. As shown in the table, VRSA delivered the best performance on both motion types.</p><p>3) Dependency on Train/Test Proportion: In order to study the degree on the performance of the model, we measured the mean values of the PLCCs over 2000 trials as a function of the training set percentage as it ranged from 10% to 90% in 10% increments. Fig. <ref type="figure">14</ref> shows as the percentage of the training set increased from 50% to 90%, the performance difference varied less than 10%.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>4) Feature Ablation Study:</head><p>We also studied the performances of the feature groups and combinations of feature groups. Table <ref type="table">IV</ref> shows the LCC and SROCC for each feature group of combinationsacross 2000 train-test trials. Specifically, we tested SF, STF, RF, EMF, and combinations (SF+STF), (SF+RF), (SF+EMF), (SF+STF+RF) and (SF+RF+EMF).   Since the spatial feature SF is not a major contributor to VR sickness, its correlation against MOS is only about 0.3. However, the frame difference map-based statistical features (STF, RF and EMF) delivered reliable predictive performance. In particular, the correlation obtained by RF were higher than those of the other single features. When SF was combined with STF, RF and EMF, the correlation SF+RF+EMF delivered even higher SROCC performance, indicating that EMF efficiently contributes to performance as well. Overall, when all the features were used together, VRSA much better performance than the other models.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>D. Content Dependence</head><p>We also explored the way that VRSA behaves across diverse video contents. Fig. <ref type="figure">15</ref> shows the performance of VRSA over the 2000 trials on each of the reference VR videos in the VR-SP database. Although the correlation it attained was slightly better on some scenarios than the others, the variation in performance lay within a fairly small range.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>E. Effect of Model Parameters</head><p>To study model behavior against the parameters of VRSA, we tabulated performance while varying the filter window sizes. Table <ref type="table">VII</ref> shows performances for four filter windows on the VR-SP database. It can be seen that the algorithm performed best for a Gaussian filter window of size 15.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>F. VRPA Prediction Results</head><p>To evaluate the performance of VRPA, we used the subjective presence assessment in Section III and the additional features defined in Section IV. Similar to the analysis of VRSA, SROCC and PLCC were used to measure performance. The overall experimental protocol was as described in Section VI-A.  As discussed in Section V, the VR sickness feature is a good prediction of the sense of VR presence, and it delivers a reliable correlation of &#8764;0.69. When all of the features are used together, VRPA attains a top correlation performance about &#8764;0.75. These results infer the necessity of considering all of the sickness, visual activity and content features simultaneously to evaluate VR presence accurately.</p><p>2) Predicted Results: To determine whether VR sickness is related to presence, four sampled VR sequences and their predicted sickness/presence scores (MOS scaled to [0, 1]) are illustrated in Fig. <ref type="figure">16</ref>. Overall, it can be seen that the predicted scores are close to MOS. Also, the sequences with high VR sickness scores, such as shown in Figs. <ref type="figure">16 (a</ref>) and (b), have relatively low VR presence scores. Similarly, as shown in Fig. <ref type="figure">16</ref> (d), when there is a very low level of VR sickness, VR presence becomes quite reduced. However, in Fig. <ref type="figure">16 (c</ref>) exemplifying an appropriate level of VR sickness, the sense of presence is relatively maximized. Broadly, it may be seen that there is a strong relationship between VR sickness and VR presence, and our prediction models agree with that observation.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>VII. CONCLUSION</head><p>Recently, virtual experience technology has improved remarkably, raising the importance of VR sickness/presence predictors in commercial HMD products. Reliable QoE assessment algorithms could help provide users with more realistic and comfortable visual experiences when using VR. In the VR industry, significant efforts have been made to improve software/hardware techniques to better deliver viewers' visual satisfaction. Here, we described a new large-scale database of VR sickness/presence and that we created, and used it to formulate new VR prediction models. Our simulation results show that our proposed models can autonomously and effectively predict VR sickness/presence, achieving much better performance than conventional algorithms. We expect that the proposed models could be applied in a variety of applications, such as VR content creation, VR broadcasting, and VR gaming.</p></div><note xmlns="http://www.tei-c.org/ns/1.0" place="foot" xml:id="foot_0"><p>Authorized licensed use limited to: University of Texas at Austin. Downloaded on June 19,2021 at 00:44:35 UTC from IEEE Xplore. Restrictions apply.</p></note>
		</body>
		</text>
</TEI>
