Abstract
Matching face images across different modalities is a challenging open problem for various reasons, notably feature heterogeneity, and particularly in the case of sketch recognition – abstraction, exaggeration and distortion. Existing studies have attempted to address this task by engineering invariant features, or learning a common subspace between the modalities. In this paper, we take a different approach and explore learning a mid-level representation within each domain that allows faces in each modality to be compared in a domain invariant way. In particular, we investigate sketch-photo face matching and go beyond the well-studied viewed sketches to tackle forensic sketches and caricatures where representations are often symbolic. We approach this by learning a facial attribute model independently in each domain that represents faces in terms of semantic properties. This representation is thus more invariant to heterogeneity, distortions and robust to mis-alignment. Our intermediate level attribute representation is then integrated synergistically with the original low-level features using CCA. Our framework shows impressive results on cross-modal matching tasks using forensic sketches, and even more challenging caricature sketches. Furthermore, we create a new dataset with ≈59, 000 attribute annotations for evaluation and to facilitate future research.
Cite
CITATION STYLE
Ouyang, S., Hospedales, T., Song, Y. Z., & Li, X. (2015). Cross-Modal face matching: Beyond viewed sketches. In Lecture Notes in Computer Science (Vol. 9004, pp. 210–225). Springer Verlag. https://doi.org/10.1007/978-3-319-16808-1_15
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.