How Much Does Each Modality Matter? A Cross-Study Empirical Comparison of Visual and Textual Signal Contributions in Multimodal Recommendation on E-Commerce Data
Downloads
Multimodal recommenders routinely report accuracy gains over interaction-only models, yet the field rarely quantifies how much of that gain is attributable to the visual channel and how much to the textual channel. This paper reports a cross-study empirical comparison built on 54 published result cells drawn from three peer-reviewed studies that share an identical evaluation protocol on three Amazon categories: the same 5-core splits, the same released 4096-dimensional visual and 384-dimensional textual features, the same 8:1:1 per-user partition, and the same all-ranking evaluation. No new architecture is proposed. Three findings emerge. The apparent value of multimodal content depends heavily on the interaction-only reference point: 40 of 42 multimodal cells improve on a matrix-factorisation baseline by 65.91% on average, while only 30 of 42 improve on a tuned LightGCN baseline, with a mean gain of 10.29%. A published per-modality ablation shows that one content channel captures most of the available benefit, the second channel adding 3.44% on average. Published conclusions about which channel dominates conflict on identical data. Modality contribution is best read as backbone-conditional rather than as a property of the data.
He, R., & McAuley, J. (2016). VBPR: Visual Bayesian personalized ranking from implicit feedback. Proceedings of the AAAI Conference on Artificial Intelligence, 30(1), 144–150.
Wei, Y., Wang, X., Nie, L., He, X., Hong, R., & Chua, T.-S. (2019). MMGCN: Multi-modal graph convolution network for personalized recommendation of micro-video. Proceedings of the 27th ACM International Conference on Multimedia, 1437–1445.
Liu, Q., Hu, J., Xiao, Y., Zhao, X., Gao, J., Wang, W., Li, Q., & Tang, J. (2024). Multimodal recommender systems: A survey. ACM Computing Surveys, 57(2), 1–17.
Zhou, H., Zhang, Y., Sun, A., & Shen, Z. (2025). Does multimodality improve recommender systems as expected? A critical analysis and future directions. arXiv preprint arXiv:2508.05377.
Yuan, Z., Yuan, F., Song, Y., Li, Y., Fu, J., Yang, F., Pan, Y., & Ni, Y. (2023). Where to go next for recommender systems? ID- vs. modality-based recommender models revisited. Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2639–2649.
Wei, Y., Wang, X., Nie, L., He, X., & Chua, T.-S. (2020). Graph-refined convolutional network for multimedia recommendation with implicit feedback. Proceedings of the 28th ACM International Conference on Multimedia, 3541–3549.
Zhang, J., Zhu, Y., Liu, Q., Wu, S., Wang, S., & Wang, L. (2021). Mining latent structures for multimedia recommendation. Proceedings of the 29th ACM International Conference on Multimedia, 3872–3880.
Zhang, J., Zhu, Y., Liu, Q., Zhang, M., Wu, S., & Wang, L. (2023). Latent structure mining with contrastive modality fusion for multimedia recommendation. IEEE Transactions on Knowledge and Data Engineering, 35(9), 9154–9167.
Zhou, X., & Shen, Z. (2023). A tale of two graphs: Freezing and denoising graph structures for multimodal recommendation. Proceedings of the 31st ACM International Conference on Multimedia, 935–943.
Yu, P., Tan, Z., Lu, G., & Bao, B.-K. (2023). Multi-view graph convolutional network for multimedia recommendation. Proceedings of the 31st ACM International Conference on Multimedia, 6576–6585.
Tao, Z., Liu, X., Xia, Y., Wang, X., Yang, L., Huang, X., & Chua, T.-S. (2023). Self-supervised learning for multimedia recommendation. IEEE Transactions on Multimedia, 25, 5107–5116.
Zhou, X., Zhou, H., Liu, Y., Zeng, Z., Miao, C., Wang, P., You, Y., & Jiang, F. (2023). Bootstrap latent representations for multi-modal recommendation. Proceedings of the ACM Web Conference 2023, 845–854.
Ferrari Dacrema, M., Cremonesi, P., & Jannach, D. (2019). Are we really making much progress? A worrying analysis of recent neural recommendation approaches. Proceedings of the 13th ACM Conference on Recommender Systems, 101–109.
Ferrari Dacrema, M., Boglio, S., Cremonesi, P., & Jannach, D. (2021). A troubling analysis of reproducibility and progress in recommender systems research. ACM Transactions on Information Systems, 39(2), 1–49.
Anelli, V. W., Bellogín, A., Ferrara, A., Malitesta, D., Merra, F. A., Pomo, C., Donini, F. M., & Di Noia, T. (2021). Elliot: A comprehensive and rigorous framework for reproducible recommender systems evaluation. Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2405–2414.
Zhou, X. (2023). MMRec: Simplifying multimodal recommendation. Proceedings of the 5th ACM International Conference on Multimedia in Asia Workshops, 1–2.
Malitesta, D., Gassi, G., Pomo, C., & Di Noia, T. (2023). Ducho: A unified framework for the extraction of multimodal features in recommendation. Proceedings of the 31st ACM International Conference on Multimedia, 9668–9671.
Wang, Q., Wei, Y., Yin, J., Wu, J., Song, X., & Nie, L. (2023). DualGNN: Dual graph neural network for multimedia recommendation. IEEE Transactions on Multimedia, 25, 1074–1084.
Guo, Z., Li, J., Li, G., Wang, C., Shi, S., & Ruan, B. (2024). LGMRec: Local and global graph learning for multimodal recommendation. Proceedings of the AAAI Conference on Artificial Intelligence, 38(8), 8454–8462.
Ni, Y., Cheng, Y., Liu, X., Fu, J., Li, Y., He, X., Zhang, Y., & Yuan, F. (2025). A content-driven micro-video recommendation dataset at scale. Proceedings of the 34th ACM International Conference on Information and Knowledge Management.
Zhou, S., Xu, F. F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Ou, T., Bisk, Y., Fried, D., Alon, U., & Neubig, G. (2024). WebArena: A realistic web environment for building autonomous agents. International Conference on Learning Representations (ICLR 2024).
Xie, T., Zhang, D., Chen, J., Li, X., Zhao, S., Cao, R., Hua, T. J., Cheng, Z., Shin, D., Lei, F., Liu, Y., Xu, Y., Zhou, S., Savarese, S., Xiong, C., Zhong, V., & Yu, T. (2024). OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments. Advances in Neural Information Processing Systems 37 (NeurIPS 2024), Datasets and Benchmarks Track.
Liu, X., Yu, H., Zhang, H., Xu, Y., Lei, X., Lai, H., Gu, Y., Ding, H., Men, K., Yang, K., Zhang, S., Deng, X., Zeng, A., Du, Z., Zhang, C., Shen, S., Zhang, T., Su, Y., Sun, H., ... Tang, J. (2024). AgentBench: Evaluating LLMs as agents. International Conference on Learning Representations (ICLR 2024).
Shang, W., Xu, W., & Zhou, Y. (2026). Feature Weight Optimization in Machine Learning Classifiers for Conflict Escalation Early Warning: Evidence from Diplomatic Signals and News Text. Journal of Sustainability, Policy, and Practice, 2(3), 26–37.
Shang, W., Zhao, F., & Liu, M. (2026). A Comparative Study of Spatiotemporal Clustering and Classification Approaches for Security Incident Risk Assessment in UN Peacekeeping Operations. Journal of Sustainability, Policy, and Practice, 2(3), 1–11.
Chen, Y., & Tang, T. (2026). Evaluating Prompt Engineering Strategies for Few-Shot Cyber Threat Intelligence Entity and Relation Extraction from Multi-Source Reports. Journal of Science, Innovation & Social Impact, 2(2), 153–164.
Li, Y., & Ling, Z. (2026). Real-Time Multi-Risk Early Warning for Community Banks: An Application of Ensemble Anomaly Detection and Explainable Artificial Intelligence. Journal of Advanced Computing Systems, 6(2), 15–27.
Li, Y., & Zhang, S. (2025). Machine Learning-Based Credit Risk Early Warning System for Small and Medium-Sized Financial Institutions: An Ensemble Learning Approach with Interpretable Risk Indicators. Journal of Science, Innovation & Social Impact, 1(1), 372–383.
Li, Y. (2026). Enhancing Financial Compliance Transparency through Automated Data Governance and Intelligent Risk Reporting. Journal of Science, Innovation & Social Impact, 2(1), 299–313.
Li, Y., Zhang, X., & Hao, C. (2026). Graph Neural Network-Based Cross-Market Risk Contagion Analysis between US Equity Sectors and Treasury Yields during the 2023 Regional Banking Stress. Journal of Sustainability, Policy, and Practice, 2(4), 154–168.
Li, Y., & Fu, X. (2026). Comparative Evaluation of Graph Neural Networks for Cross-Market Risk Contagion Path Identification in Multi-Layer Financial Networks. Journal of Sustainability, Policy, and Practice, 2(3), 1–14.
Li, Y., Zhao, F., & Hu, J. (2026). Identifying Cross-Market Risk Contagion Amplifiers via Graph Attention Networks: Empirical Evidence from US Financial Stress Periods. Journal of Computing Innovations and Applications, 4(1), 164–175.
Chen, Y., & Hu, J. (2026). Graph Neural Network-Based Cascading Disruption Path Identification in Multi-Tier Rare Earth Processing Networks. Journal of Global Engineering Review, 4(1), 99–112.
Chen, Y., & Lai, J. (2026). Multi-Metric Trustworthiness Evaluation of AI-Assisted Medical Imaging Diagnosis: Integrating Confidence Calibration and Distribution Shift Detection. Journal of Global Engineering Review, 4(1), 113–126.
Chen, Y. (2026). Multi-Dimensional Training Data Bias Detection and Fairness-Aware Augmentation for Equitable AI-Assisted Medical Imaging Diagnosis. Journal of Sustainability, Policy, and Practice, 2(4), 1–15.
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2023). ReAct: Synergizing reasoning and acting in language models. International Conference on Learning Representations (ICLR 2023).
Shinn, N., Cassano, F., Berman, E., Gopinath, A., Narasimhan, K., & Yao, S. (2023). Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems 36 (NeurIPS 2023).
Wang, L., Xu, W., Lan, Y., Hu, Z., Lan, Y., Lee, R. K.-W., & Lim, E.-P. (2023). Plan-and-Solve prompting: Improving zero-shot chain-of-thought reasoning by large language models. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL 2023), 2609–2634.
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., & Narasimhan, K. (2023). Tree of Thoughts: Deliberate problem solving with large language models. Advances in Neural Information Processing Systems 36 (NeurIPS 2023).
Chen, Y. (2024). Explainable Attack Path Reasoning for Industrial Control Network Security Based on Knowledge Graphs. Journal of Computing Innovations and Applications, 2(1), 128–139.
Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., & Anandkumar, A. (2024). Voyager: An open-ended embodied agent with large language models. Transactions on Machine Learning Research.
Zheng, B., Gou, B., Kil, J., Sun, H., & Su, Y. (2024). GPT-4V(ision) is a generalist web agent, if grounded. Proceedings of the 41st International Conference on Machine Learning (ICML 2024).
He, H., Yao, W., Ma, K., Yu, W., Dai, Y., Zhang, H., Lan, Z., & Yu, D. (2024). WebVoyager: Building an end-to-end web agent with large multimodal models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024).
Cao, H., Hu, J., & Luo, C. (2025). Behavioural Feature Analysis for Anomalous Click Detection in Mobile Advertising Environments: Toward In-App Browser-Specific Detection. Journal of Computing Innovations and Applications, 3(2), 96–105.
Cao, H. (2026). An Empirical Analysis of Click Temporal Features for Automated Ad Fraud Detection in Mobile In-App Browser Environments. Journal of Sustainability, Policy, and Practice, 2(4), 142–153.
Liu, E. Z., Guu, K., Pasupat, P., Shi, T., & Liang, P. (2018). Reinforcement learning on web interfaces using workflow-guided exploration. International Conference on Learning Representations (ICLR 2018).
Mialon, G., Fourrier, C., Swift, C., Wolf, T., LeCun, Y., & Scialom, T. (2024). GAIA: A benchmark for general AI assistants. International Conference on Learning Representations (ICLR 2024).
Peng, Y., & Gao, S. (2026). An Empirical Comparison of XGBoost and LightGBM for Capital Expenditure Deviation Prediction in US FERC-Regulated Rate-Base Transmission Projects. Journal of Sustainability, Policy, and Practice, 2(4), 90–102.
Hu, J., Wang, X., & Lai, J. (2026). Benchmarking Learned Cardinality Estimation Techniques for Analytical Query Processing in Data Warehouses. Journal of Computer Technology and Applied Mathematics, 3(3), 1–8.
Hu, J., & Long, X. (2024). Graph Learning-Based Behavioral Detection for Software Supply Chain Attacks. Journal of Advanced Computing Systems, 4(4), 49–60.
Wen, S., & Tang, T. (2025). A Comparative Evaluation of URL-Sharing, Content Similarity, and Temporal Synchronicity Signals for Detecting Coordinated Inauthentic Behavior in Multilingual Political Discourse. Journal of Global Engineering Review, 3(2), 69–78.
Koh, J. Y., Lo, R., Jang, L., Duvvur, V., Lim, M. C., Huang, P.-Y., Neubig, G., Zhou, S., Salakhutdinov, R., & Fried, D. (2024). VisualWebArena: Evaluating multimodal agents on realistic visual web tasks. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024), 881–905.
Deng, X., Gu, Y., Zheng, B., Chen, S., Stevens, S., Wang, B., Sun, H., & Su, Y. (2023). Mind2Web: Towards a generalist agent for the web. Advances in Neural Information Processing Systems 36 (NeurIPS 2023), Datasets and Benchmarks Track.
Yao, S., Chen, H., Yang, J., & Narasimhan, K. (2022). WebShop: Towards scalable real-world web interaction with grounded language agents. Advances in Neural Information Processing Systems 35 (NeurIPS 2022).
Cao, H., & Shi, W. (2026). Statistical Anomaly Detection Approach for Field Mapping Validation in Enterprise Payroll Data Migration. Journal of Computing Innovations and Applications, 4(1), 137–153.
Cao, H., & Liu, M. (2026). Empirical Evaluation of Constraint Discovery Techniques for Business Logic Consistency Verification in Payroll Data Migration. Journal of Sustainability, Policy, and Practice, 2(3), 92–103.
Cao, H., & Long, L. (2026). Empirical Evaluation of Multi-Source Monitoring Signal Effectiveness and Lead Time for Performance Degradation Prediction in Kubernetes-Based Microservices. Journal of Advanced Computing Systems, 6(4), 15–26.
Long, X., Hu, J., & Ling, Z. (2026). A Comparative Analysis of Telemetry-Driven Anomaly Detection Approaches for Dual-Purpose Operational and Security Optimization in Edge Computing Infrastructure. Journal of Computing Innovations and Applications, 4(1), 79–88.
Cao, H. (2024). Privacy-Preserving Click Pattern Anomaly Detection for Mobile In-App Browser Advertising Fraud. Journal of Computing Innovations and Applications, 2(2), 151–161.
Cao, H. (2024). Detecting Fraudulent Click Patterns in Mobile In-App Browsers: A Multi-dimensional Behavioral Analysis Approach. Artificial Intelligence and Machine Learning Review, 5(2), 130–142.
Qin, Y., Liang, S., Ye, Y., Zhu, K., Yan, L., Lu, Y., Lin, Y., Cong, X., Tang, X., Qian, B., Zhao, S., Hong, L., Tian, R., Xie, R., Zhou, J., Gerstein, M., Li, D., Liu, Z., & Sun, M. (2024). ToolLLM: Facilitating large language models to master 16000+ real-world APIs. International Conference on Learning Representations (ICLR 2024).
Cemri, M., Pan, M. Z., Yang, S., Agrawal, L., Chopra, B., Tiwari, R., Keutzer, K., Parameswaran, A., Klein, D., Ramchandran, K., Gonzalez, J. E., Zaharia, M., & Stoica, I. (2025). Why do multi-agent LLM systems fail? arXiv preprint arXiv:2503.13657.
Chen, Y., & Chen, Z. (2025). Multi-Objective Deep Reinforcement Learning for Carbon-Aware Spatiotemporal Workload Scheduling in Geo-Distributed Data Centers. Journal of Advanced Computing Systems, 5(10), 18–30.
Long, L., & Hu, J. (2026). Multi-Objective Particle Swarm Optimization for Site Selection and Policy Subsidy Maximization of Foreign Renewable Energy Enterprises in the United States. Artificial Intelligence and Machine Learning Review, 7(2), 54–69.
Li, Y., & Long, L. (2026). Lightweight AI-Driven Stress Testing for Small and Medium Financial Institutions: A Variational Autoencoder Approach with Extreme Value Theory for Macroeconomic Scenario Generation. Artificial Intelligence and Machine Learning Review, 7(1), 108–119.
Wu, C., Guan, H., & Weng, H. (2024). Forecasting Hospital Resource Demand Using Gradient Boosting: An Operational Analytics Approach for Bed Allocation and Patient Flow Management. Journal of Computing Innovations and Applications, 2(1), 74–85.
Chen, Y., Chen, Z., & Zou, D. (2025). CarbonShift: Harnessing Grid Carbon Variability for Geo-Distributed Workload Scheduling. Artificial Intelligence and Machine Learning Review, 6(4), 18–31.
Wei, C., & Wu, C. (2024). Credit Risk Transmission Mechanism and Prevention Strategies in Supply Chain Finance: A Core Enterprise Perspective. Artificial Intelligence and Machine Learning Review, 5(2), 101–115.
Shridhar, M., Yuan, X., Côté, M.-A., Bisk, Y., Trischler, A., & Hausknecht, M. (2021). ALFWorld: Aligning text and embodied environments for interactive learning. International Conference on Learning Representations (ICLR 2021).
Patil, S. G., Zhang, T., Wang, X., & Gonzalez, J. E. (2024). Gorilla: Large language model connected with massive APIs. Advances in Neural Information Processing Systems 37 (NeurIPS 2024).
Wu, C., & Pan, Z. (2024). An Integrated Graph Neural Network and Reinforcement Learning Framework for Intelligent Drug Discovery. Journal of Advanced Computing Systems, 4(6), 19–29.
Furuta, H., Lee, K.-H., Nachum, O., Matsuo, Y., Faust, A., Gu, S. S., & Gur, I. (2024). Multimodal web navigation with instruction-finetuned foundation models. International Conference on Learning Representations (ICLR 2024).
Li, Y., Wu, C., Li, L., Liu, Y., & Zhu, J. (2021). Caption generation from road images for traffic scene modeling. IEEE Transactions on Intelligent Transportation Systems, 23(7), 7805–7816.
Li, L., Li, Y., Wu, C., Dong, H., Jiang, P., & Wang, F. (2021). Detail fusion GAN: High-quality translation for unpaired images with GAN-based data augmentation. 2020 25th International Conference on Pattern Recognition (ICPR), 1731–1736.
Huo, H., Li, Y., Wu, C., Wu, X., Tian, Z., & Liu, Y. (2019). Semantic segmentation and scene reconstruction for traffic simulation using CNN. 2019 2nd China Symposium on Cognitive Computing and Hybrid Intelligence.
Wu, X., Li, Y., Hao, Z., Wu, C., Wang, X., & Liu, Y. (2019). Image style transformation based on structure GAN. 2019 Chinese Automation Congress (CAC), 2002–2007.
Wu, C., Li, Y., Li, L., Wang, L., & Liu, Y. (2020). Caption generation from road images for traffic scene construction. 2020 IEEE Intelligent Vehicles Symposium (IV), 1271–1276.
Wu, X., Li, Y., Liu, Y., Pang, S., Wang, L., Wu, C., & Huo, H. (2019). Jointly detecting and retrieving vehicles from road image sequences based on CNN. 2019 IEEE Intelligent Vehicles Symposium (IV), 2530–2535.
Shi, X., & Weng, H. (2024). Comparative analysis of unsupervised learning approaches for anomalous billing pattern detection in healthcare payment integrity. Journal of Computing Innovations and Applications, 2(1), 111–127.
Weng, H., Yang, Q., Fu, H., Pan, H., & Lv, X. (2026). When retrieval hurts code completion: A diagnostic study of stale repository context. arXiv preprint arXiv:2605.14478.
Weng, H. (2025). Deep embedding clustering with adaptive feature selection for banking customer segmentation. Spectrum of Research, 5(2).
Weng, H., & Lei, Y. (2024). Cross-modal artifact mining for generalizable deepfake detection in the wild. Journal of Computing Innovations and Applications, 2(2), 78–87.
Weng, H., & Li, X. (2024). Renewable-aware cooperative scheduling for distributed AI training across geo-distributed data centers. Artificial Intelligence and Machine Learning Review, 5(2), 91–100.
Zhang, S., Wang, Y., & Weng, H. (2024). Industrial IoT anomaly detection using improved autoencoder architecture. Artificial Intelligence and Machine Learning Review, 5(1), 67–78.
Weng, H., Wang, H., & Wei, C. (2024). Adaptive bidding strategies for hybrid auction mechanisms in programmatic advertising. Journal of Advanced Computing Systems, 4(4), 13–25.
Weng, H., Zhang, S., & Min, S. (2024). Multi-constraint optimization for real-time bidding: A reinforcement learning approach. Artificial Intelligence and Machine Learning Review, 5(1), 93–104.
Weng, H. (2026). Predictive power versus feature stability: An empirical comparison of filter, embedded, and wrapper feature selection methods for consumer credit pre-approval. Academia Nexus Journal, 5(1).
