A sub-1B parameter multimodal emotion recognition model that processes video, audio, and text to recognize emotions. Upload a video or audio file and ask a question about the emotional state.
Paper: Do We Really Need Multimodal Emotion Language Models Larger Than 1B Parameters? | GitHub | Model