<?xml-model href='http://www.tei-c.org/release/xml/tei/custom/schema/relaxng/tei_all.rng' schematypens='http://relaxng.org/ns/structure/1.0'?><TEI xmlns="http://www.tei-c.org/ns/1.0">
	<teiHeader>
		<fileDesc>
			<titleStmt><title level='a'>Learning to See People Like People: Predicting Social Impressions of Faces</title></titleStmt>
			<publicationStmt>
				<publisher></publisher>
				<date>2017</date>
			</publicationStmt>
			<sourceDesc>
				<bibl> 
					<idno type="par_id">10072403</idno>
					<idno type="doi"></idno>
					<title level='j'>Cognitive Science</title>
<idno></idno>
<biblScope unit="volume"></biblScope>
<biblScope unit="issue"></biblScope>					

					<author>A Song</author><author>L Linjie</author><author>C Atalla</author><author>G. Gottrell</author>
				</bibl>
			</sourceDesc>
		</fileDesc>
		<profileDesc>
			<abstract><ab><![CDATA[Humans make complex inferences on faces, ranging from objective properties (gender, ethnicity, expression, age, identity,  etc)  to subjective judgments (facial attractiveness, trustworthiness, sociability, friendliness, etc). While the objective aspects of face perception have been extensively studied, relatively fewer computational models have been developed for the social impressions of faces. Bridging this gap, we develop a method to predict human impressions of faces in 40 subjective social dimensions, using deep representations from state-of-the-art neural networks. We find that model performance grows as the human consensus on a face trait increases, and that model predictions outperform human groups in correlation with human averages. This illustrates the learnability of subjective social perception of faces, especially when there is high human consensus. Our system can be used to decide which photographs from a personal collection will make the best impression. The results are significant for the field of social robotics, demonstrating that robots can learn the subjective judgments defining the underlying fabric of human interaction.]]></ab></abstract>
		</profileDesc>
	</teiHeader>
	<text><body xmlns="http://www.tei-c.org/ns/1.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:xlink="http://www.w3.org/1999/xlink">
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Introduction</head><p>With the huge success of deep learning techniques, current state-of-the-art computer vision algorithms have approached or exceeded human ability in recognizing a face <ref type="bibr">(Taigman, Yang, Ranzato, &amp; Wolf, 2014;</ref><ref type="bibr">Stewart, Andriluka, &amp; Ng, 2016)</ref> and identifying the objective properties of a face, such as age and gender estimation, <ref type="bibr">(Guo, Fu, Dyer, &amp; Huang, 2008)</ref>. However, humans not only read objective properties from a face, like expression, age, and identity, but also form subjective impressions of social aspects of a face <ref type="bibr">(Todorov, Olivola, Dotsch, &amp; Mende-Siedlecki, 2015)</ref> at first sight, such as facial attractiveness <ref type="bibr">(Thornhill &amp; Gangestad, 1999)</ref>, friendliness, trustworthiness <ref type="bibr">(Todorov, Baron, &amp; Oosterhof, 2008)</ref>, sociability, dominance <ref type="bibr">(Mignault &amp; Chaudhuri, 2003)</ref>, and typicality. In spite of the subjective nature of social perceptions, there is often a consensus among human in how they perceive attractiveness, trustworthiness, and dominance &#8224; These authors contributed equally. in faces <ref type="bibr">(Falvello, Vinson, Ferrari, &amp; Todorov, 2015;</ref><ref type="bibr">Eisenthal, Dror, &amp; Ruppin, 2006)</ref>. This indicates that faces contain high-level visual cues for social inferences, therefore making it possible to model the inference process computationally. Social judgments, as an important part of people's daily interactions, have a significant impact on social outcomes, ranging from electoral success to sentencing decisions <ref type="bibr">(Oosterhof &amp; Todorov, 2008;</ref><ref type="bibr">Willis &amp; Todorov, 2006)</ref>.</p><p>Are deep learning models, which are successful in various visual tasks, also capable of predicting subjective social impressions of faces? Even before the advent of deep learning, there have been models using traditional computer vision algorithms and simulated faces to model the perception of facial attractiveness <ref type="bibr">(Thornhill &amp; Gangestad, 1999;</ref><ref type="bibr">Eisenthal et al., 2006;</ref><ref type="bibr">Kagian et al., 2008;</ref><ref type="bibr">Gray, Yu, Xu, &amp; Gong, 2010)</ref>, trustworthiness <ref type="bibr">(Falvello et al., 2015;</ref><ref type="bibr">Todorov, Baron, &amp; Oosterhof, 2008)</ref>, sociability, aggressiveness <ref type="bibr">(Mignault &amp; Chaudhuri, 2003)</ref>, familiarity <ref type="bibr">(Peskin &amp; Newell, 2004)</ref>, and memorability <ref type="bibr">(Bainbridge, Isola, &amp; Oliva, 2013;</ref><ref type="bibr">Khosla, Bainbridge, Torralba, &amp; Oliva, 2013)</ref>. Recently, there has been work on modeling the "big five " personality traits perceived by humans when viewing another person in video clips <ref type="bibr">(Escalera et al., 2016)</ref>.</p><p>In this paper, we examine human social perceptions of faces in 40 dimensions extensively and systematically. We evaluate the human consistency and correlation in 40 social features (20 relevant pairs) that are typically studied by social psychologists <ref type="bibr">(Todorov, Said, Engell, &amp; Oosterhof, 2008)</ref>, and relevant to social interactions <ref type="bibr">(Todorov et al., 2015;</ref><ref type="bibr">Oosterhof &amp; Todorov, 2008)</ref>, and use state-of-the-art deep learning algorithms to model all 40 of them. Using the internal representations learned from the deep learning models, our model can successfully predict human social perception whenever human have a consensus. We further visualize the key features defining different social attributes to facilitate a understanding of what makes a face salient in a certain social dimension.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Methods Dataset</head><p>To predict human social impressions of faces, we use a public dataset <ref type="bibr">(Bainbridge et al., 2013)</ref> consisting of 2,222 face images and annotations for 40 social attributes. Each attribute is rated on a scale of 1-9 by 15 subjects. We take the average rating from all raters as a collective estimation of human judgment for the social features of each face.</p><p>The 40 social attributes consist of 20 pairs of related traits: (attractive, unattractive), (happy, unhappy), (friendly, unfriendly), etc. Some of these traits are highly correlated and predictable from others, especially within the trait pairs. To understand the human-perceived correlations between these traits, we compute the Spearman's rank correlation between the average human ratings of every pair of social features and show their correlations in a heatmap (Figure <ref type="figure">1(a)</ref>). We order traits in the map based on similarity and positive or negative connotation. From the figure, we see that negative social features such as untrustworthy, aggressive, cold, introverted, and irresponsible form a correlated block. Likewise, the most positive features such as attractive, sociable, caring, friendly, happy, intelligent, interesting, and confident are highly correlated with each other. Although we choose 20 pairs of opposite features, they are not completely complementary and redundant. Principal Component Analysis of the covariance matrix shows that it takes 24 principal components to cover 95% of the variance.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Regression Model for Social Attributes</head><p>After averaging human ratings, each face receives a continuous score from 1 to 9 in all social dimensions. We model these social scores with a regression model. We propose a ridge regression model on either features from deep convolutional neural networks (CNN) or traditional face geometry based features, and present results from both feature sets. Such visual features are usually high-dimensional, so we first perform Principal Component Analysis (PCA) on the extracted features of the training set to reduce dimensionality. The PCA dimensionality is chosen by cross-validation on a validation set, separately for each trait. The PCA weights are saved and further used in fine-tuning our CNN-regression model.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Regression on Geometric Features</head><p>Past studies have found that facial attractiveness can be inferred from the geometric ratios and configurations of a face <ref type="bibr">(Eisenthal et al., 2006;</ref><ref type="bibr">Kagian et al., 2008)</ref>. We suggest that other social attributes can also be inferred from geometric features. We compute 29 geometric features based on definitions described in <ref type="bibr">(Ma, Correll, &amp; Wittenbrink, 2015)</ref>, and further extract a 'smoothness' feature and 'skin color' feature according to the procedure in <ref type="bibr">(Eisenthal et al., 2006;</ref><ref type="bibr">Kagian et al., 2008)</ref>. The smoothness of a face was evaluated by applying a Canny edge detector to regions from the cheek and forehead areas <ref type="bibr">(Eisenthal et al., 2006)</ref>. The more edges detected, the less smooth the skin is. The regions we chose to compute smoothness and skin color are highlighted in the right subplot of Figure <ref type="figure">2</ref>. The skin color feature is extracted from the same region as smoothness, converted from RGB to HSV. However, regressing on these handcrafted features alone is not enough to capture the richness of geometric details in a face. We therefore use a computer vision library (dlib, C++) to automatically label 68 face landmarks (see Figure <ref type="figure">2</ref>) for each face, and then compute distances and slopes between any two landmarks. Combining 29 handcrafted geometric features, smoothness, color and the distance-slope features, we obtain 4592 features in total. Since the features are highly correlated, we apply PCA to reduce dimensionality. Again, the PCA dimensionality is chosen by cross-validating on the hold out set separately for each facial attribute. Then a ridge regression model is applied to predict social attribute ratings of a face. The hyper-parameter of ridge regression is selected by leave-one-out validation within the training set.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Regression on CNN Features</head><p>Previous studies have shown that pretrained deep learning models can provide feature representations versatile for related tasks. We therefore extract image features from pretrained neural networks, choosing from six architectures with different original training goals: (1) VGG16, trained for object recognition <ref type="bibr">(Simonyan &amp; Zisserman, 2014)</ref>, (2) VGG-Face, trained for face identification <ref type="bibr">(Simonyan &amp; Zisserman, 2014)</ref>, (3) AlexNet, trained for object classification <ref type="bibr">(Krizhevsky, Sutskever, &amp; Hinton, 2012)</ref>, (4) Inception from Google, trained for object recognition <ref type="bibr">(Szegedy et al., 2015)</ref>, (5) a shallow Siamese neural network that we train from scratch to cluster faces by identity, (6) a state of the art VGG-derived network (Face-LandmarkNN) trained for the face landmark localization task.</p><p>To find the best CNN features among the six networks, we first find the best-performing feature layers of each network in the ridge regression prediction task. Before the ridge regression, we perform PCA and pick the PCA dimensionality that gives best results on the validation set. Then, we compare the results among networks to select the best features overall.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Results</head><p>After comparing all 6 networks, we find that the conv5 2 layer of VGG16 (trained for object classification) lead to the best results. This set of features significantly outperforms the three networks trained solely on faces, while also slightly outperforming AlexNet and Inception networks. These bestperforming CNN features also exceed the prediction correlation of the geometric features in most attributes. Figure <ref type="figure">3</ref> compares prediction performance of the CNN model and the geometric feature model.</p><p>We speculate that the poor performance from the face recognition networks can be attributed to their optimization for specific facial tasks. Learning face landmark configurations and differences between faces that define identity may not correlate well with the task at hand, which looks for commonalities behind certain social features beyond identity. The consider them in a decision. A robot need not treat an attractive or unattractive person differently for its own purposes, but this knowledge could affect how interactions are made for the sake of the human, knowing in advance how that person may feel that they fit into the social landscape.</p><p>Expansions on this work may include investigating the image properties that determine high level social features, beyond the attractiveness features we display in Figure <ref type="figure">5</ref>. Additionally, social trait prediction may benefit from a single model with a shared representation, while this paper approaches each attribute as a separate regression task.</p><p>For future work, we aim to develop a generative model which can automatically modify a face's attributes (either objective or subjective) while preserving its realism and identity. Practically speaking, such a model could improve a face's perceived social features in positive ways (e.g. make a face look more sociable, trustworthy). More importantly, it would enable psychologists to quantify human biases during the formation of social impression in a precise and systematic manner. Psychologists could generate variants of a real face differing in age, gender, race, and explore how various factors separately and jointly affect the social impressions of faces.</p></div>		</body>
		</text>
</TEI>
