Multimodal Test Item Parameter Prediction From Text, Images, and Metadata: Fusing Together AI Vision and Language Models
Educational and Psychological Measurement
Published online on July 13, 2026
Abstract
Educational and Psychological Measurement, Ahead of Print.
We propose a flexible multimodal model for predicting all dichotomous and polytomous item parameters from text, images, and metadata by fusing representations from encoder Transformer vision and language models. This deep learning model accommodates ...
We propose a flexible multimodal model for predicting all dichotomous and polytomous item parameters from text, images, and metadata by fusing representations from encoder Transformer vision and language models. This deep learning model accommodates ...