Vision Transformer with Local Binary Pattern: A Contrastive Learning-Based Local–Global Hybrid Model for Facial Expression Recognition
Main Article Content
Abstract
Facial expression recognition is a critical component in affective computing and human-computer interaction. Automated facial expression recognition systems have shown impressive results lately due to Deep Learning innovated techniques. On the other hand, handcrafted methods still offer a good ability to extract micro patterns. This paper proposes a novel hybrid framework that integrates Local Binary Pattern (LBP) texture features with a Vision Transformer (ViT) backbone as one of the most effective deep learning approachs in computer vision; optimized using a joint loss function combining cross-entropy and supervised contrastive learning. The LBP features enhance local texture representation, while the ViT captures global spatial dependencies. The supervised contrastive loss promotes discriminative feature embeddings by maximizing inter-class separation and minimizing intra-class variance. Experimental evaluation on two benchmark datasets namely CK+ and JAFFE, demonstrates the superiority of the proposed method, achieving 97.45% and 95.24% accuracy respectively. These results highlight the effectiveness and generalizability of combining handcrafted and deep features in automated facial expression recognition tasks.