How Much Does Each Modality Matter? A Cross-Study Empirical Comparison of Visual and Textual Signal Contributions in Multimodal Recommendation on E-Commerce Data

multimodal recommendation; modality contribution; reproducibility; empirical comparison

Authors

  • Redwood Meridian College School of Computing and Information Sciences | Undergraduate Program in Computer Science | Portland, Oregon, USA
August 11, 2026
August 31, 2026

Downloads

Multimodal recommenders routinely report accuracy gains over interaction-only models, yet the field rarely quantifies how much of that gain is attributable to the visual channel and how much to the textual channel. This paper reports a cross-study empirical comparison built on 54 published result cells drawn from three peer-reviewed studies that share an identical evaluation protocol on three Amazon categories: the same 5-core splits, the same released 4096-dimensional visual and 384-dimensional textual features, the same 8:1:1 per-user partition, and the same all-ranking evaluation. No new architecture is proposed. Three findings emerge. The apparent value of multimodal content depends heavily on the interaction-only reference point: 40 of 42 multimodal cells improve on a matrix-factorisation baseline by 65.91% on average, while only 30 of 42 improve on a tuned LightGCN baseline, with a mean gain of 10.29%. A published per-modality ablation shows that one content channel captures most of the available benefit, the second channel adding 3.44% on average. Published conclusions about which channel dominates conflict on identical data. Modality contribution is best read as backbone-conditional rather than as a property of the data.