| Track | Chair | Venue | Time | Note |
|---|---|---|---|---|
| Workshop 1 | Luntian Mou | Ballroom | 09:00-10:00, Sun. 9 Aug. | 4 Papers |
| Workshop 2 | Wanwan Bu | Komes 2 | 09:00-12:30, Sun. 9 Aug. | https://www.ipmml.org/ |
| Tutorial | Iole Moccagatta | Komes 1 | 10:30-12:30, Sun. 9 Aug. | |
| Special Session 1 | Song Wang | Komes 1 | 13:30-14:00, Sun. 9 Aug. | 2 Papers |
| Special Session 2 | Teng-Wen Chang | Ballroom | 13:30-15:15, Sun. 9 Aug. | 7 Papers |
| Special Session 3 | Lei Gao | Komes 1 | 14:00-15:00, Sun. 9 Aug. | 4 Papers |
| Opening | Iole Moccagatta | Ballroom | 15:45-16:15, Sun. 9 Aug. | |
| Industry Talk Session | Yukiko Yanagawa | Ballroom | 16:20-17:50, Sun. 9 Aug. | Chenyang Lyu Chiraphat Boonnag Shiv Kumar Saini |
| Distinguished Plenary Talk | Paisarn Muneesawang | Ballroom | 18:00-18:45, Sun. 9 Aug. | |
| Demo Session | Kim-Hui Yap | Komes 2 | 13:30-15:15/16:15-18:00, Sun. 9 Aug. |
Chair: Luntian Mou | Venue: Ballroom | Time: 09:00-10:00, Sunday, 9 August 2026
| Time | Slot | Paper ID | Paper Title | Authors |
|---|---|---|---|---|
| 09:00-09:15 | WS-1-1 | 135 | Explainable Fundus-Based AI Glaucoma Screening: Design and Evaluation of a Healthcare System | Allan Pascual Erasmo (Yoobee Colleges); Jaya Tiwari (Waikato Hospital); Samaneh Chaharmiri Dokhaharani (Mae Fah Luang University); Mohammad Norouzifard (The University of Auckland)* |
| 09:15-09:30 | WS-1-2 | 162 | RadLens: Radiology Report Generation Using a Vision Transformer and Retrieval-Augmented Generation | Philip Niron Nithianandan (Yoobee College of Creative Innovation); Jaya Tiwari (Waikato Hospital); Seyedhashem Davarpanah (Agentic AI Hub); Mohammad Norouzifard (The University of Auckland)* |
| 09:30-09:45 | WS-1-3 | 171 | Ethical Implications of AI-Assisted Digital Phenotyping in Clinical Psychiatry | Deepa Shukla (IITJ)*; Nisha Daga (IITJ); George Kodimattam Joseph (IITJ) |
| 09:45-10:00 | WS-1-4 | 172 | Neuro-Adaptive Virtual Reality for Spider Phobia: Integrating EEG and Heart Rate Biofeedback in Closed-Loop Exposure Therapy | Rahul Kumar Ray (FLAME University); Rudra Krishna (FLAME University); Chiranjoy Chattopadhyay (FLAME University)* |
| Time | Title | Speaker | Venue |
|---|---|---|---|
| 10:30-12:30, Sun. 9 Aug. | Design and Optimization for Interactive Spatial-Temporal Multi-Modal Gaussian Splatting | Dr. Guan-Ming Su (Dolby Laboratories, USA) |
Komes 1 |
Chair: Song Wang | Venue: Komes 1 | Time: 13:30-14:00, Sunday, 9 August 2026
| Time | Slot | Paper ID | Paper Title | Authors |
|---|---|---|---|---|
| 13:30-13:45 | SS-1-1 | 124 | Arbiter: Unifying Executable Verification and Semantic Assessment | Yuhang Lu (Zhengzhou university)*; Yong Huang (Zhengzhou university); Ying Yu (The PLA Information Engineering University); Lei Li (The PLA Information Engineering University); Tao Yang (Zhengzhou university); Yun Tie (Zhengzhou university) |
| 13:45-14:00 | SS-1-2 | 129 | A Diffusion-Based Data Augmentation Framework for Semantic Segmentation with Limited Data | Zhiyan Wang (Henan University of Technology)*; Xinyu Li (Zhengzhou University); Chunhua Zhu (Henan University of Technology); Lin Qi (Zhengzhou University) |
Chair: Teng-Wen Chang | Venue: Ballroom | Time: 13:30-15:15, Sunday, 9 August 2026
| Time | Slot | Paper ID | Paper Title | Authors |
|---|---|---|---|---|
| 13:30-13:45 | SS-2-1 | 131 | The Domain-Specific Fidelity Gap: Trajectory-Based T2I Alignment in AI-Mediated Cross-Disciplinary Design Collaboration | Chen-Syuan Lin (National Yunlin University of Science and Technology); Teng-Wen Chang (National Yunlin University of Science and Technology)* |
| 13:45-14:00 | SS-2-2 | 134 | How Do Ideas Evolve Across HumanAI Co-Ideation Cycles? A Trajectory Analysis of Conversational AI in a Semester-Long Interactive-Installation Design Studio | Chor-Kheng Lim (Yuan Ze University)* |
| 14:00-14:15 | SS-2-3 | 128 | Surfacing Stress Subtypes through Cognitive-Variation-Conditioned Agentic LLM Co-Ideation: A Walkthrough Probe of Behavioral Trajectories | Siao-Rou Jheng (National Yunlin University of Science and Technology); Ya-Chen Chang (National Yunlin University of Science and Technology); Teng-Wen Chang (National Yunlin University of Science and Technology)* |
| 14:15-14:30 | SS-2-4 | 125 | Beyond Semantic Distance: Fusing Physiological and Linguistic Signals for Multimodal Coordination Assessment in Cross-Disciplinary Design | Teng-Wen Chang (National Yunlin University of Science and Technology)*; Yi-Sheng (Ezra) Wu (National Yunlin University of Science and Technology); Ya-Chen Chang (National Yunlin University of Science and Technology); Jia-Rong (Lotus) Li (National Yunlin University of Science and Technology); Siao-Rou Jheng (National Yunlin University of Science and Technology); Wu-Haw Jue (National Formosa University) |
| 14:30-14:45 | SS-2-5 | 122 | Toward AI-Mediated Attention Regulation: A Stability-Oriented Closed-Loop Architecture Using Consumer VR IMU as Wearable Behavioral Sensor | Chia Hui Nico Lo (YuanZe University)* |
| 14:45-15:00 | SS-2-6 | 127 | Structured Literary Knowledge as AI-Mediated Generative Mediation: From Cultural Interpretation to Immersive Scene Construction | Tse Wei Hsu (YunTech)* |
| 15:00-15:15 | SS-2-7 | 123 | Creatura Era: Multimodal Generative AI and Tangible AR for Cross-Disciplinary Design Collaboration | Wingly Shih (NYUST)*; Kathryn Su (NYUST) |
Chair: Lei Gao | Venue: Komes 1 | Time: 14:00-15:00, Sunday, 9 August 2026
| Time | Slot | Paper ID | Paper Title | Authors |
|---|---|---|---|---|
| 14:00-14:15 | SS-3-1 | 113 | Capturing Dynamics: The 4D Facial Expression Intensity Dataset | Zesheng Wang (University of Nantes)*; Alexandre Bruckert (University of Nantes); Pierre Lebreton (University of Nantes); Patrick Le Callet (University of Nantes); Yante Li (University of Oulu); Guoying Zhao (University of Oulu) |
| 14:15-14:30 | SS-3-2 | 170 | EmotionBench: Benchmarking Facial Emotion Understanding in Multimodal Large Language Models | Linhan Xu (Shanghai Jiao Tong University); Baiyi Chen ( Shanghai Jiao Tong University); Yixuan Gao (Shanghai Jiao Tong University)*; Huiyu Duan ( Shanghai Jiao Tong University); Yuqin Cao (Shanghai Jiao Tong University); Sijing Wu (Shanghai Jiao Tong University); Yunhao Li (Shanghai Jiao Tong University); Xiongkuo Min (Shanghai Jiao Tong University); Guangtao Zhai (Shanghai Jiao Tong University)" |
| 14:30-14:45 | SS-3-3 | 140 | Towards Emotion-Aware Wearables: Distilling PPG Foundation Models for Real-Time Affective and Physiological Monitoring | Duc Hoang Long Nguyen (University of Engineering and Technology, Vietnam National Universit)*; Patrick Le Callet (Nantes Université, Ecole Centrale Nantes); Thanh Ha Le (University of Engineering and Technology, Vietnam National Universit); Thi Duyen Ngo (University of Engineering and Technology, Vietnam National Universit) |
| 14:45-15:00 | SS-3-4 | 159 | Involuntary Leakage, Epistemic Harm, and Regulatory Arbitrage: A Cross-Jurisdictional Analysis of Micro-Expression Governance | Jenjira Jaimunk (Department of Computer Engineering, Faculty of Engineering, Chiang Mai University)*; Patrick Le Callet (Équipe Image, Perception et Interaction, Laboratoire des Sciences du Numérique de Nantes, Université de Nantes); Pradorn Sureephong (College of Arts, Media and Technology, Chiang Mai University) |
| Time | Chair | Title | Speaker | Venue |
|---|---|---|---|---|
| 18:00-18:45, Sun. 9 Aug. | Paisarn Muneesawang | Learning beyond Deep Learning: Mathematics-Inspired Models for Multi-View Analysis | Prof. Ling Guan (Toronto Metropolitan University, Canada) |
Ballroom |
Chair: Kim-Hui Yap | Venue: Komes 2 | Time: 13:30-15:15/16:15-18:00, Sunday, 9 August 2026
| Time | Slot | Paper ID | Paper Title | Authors |
|---|---|---|---|---|
| 13:30-15:15 16:15-18:00 |
DS-1-1 | 110 | Multi Agent LLMs for Diversified Large Scale Spatial Consistent 3D Scene Data Synthesis | Sandeep Bangalore (Dolby Labs)*; Lakshmi Priya Muraleedharan (Dolby Labs); Arpitha Nagaraj Hegde (Dolby Labs); Guan-Ming Su (Dolby Labs) |
| 13:30-15:15 16:15-18:00 |
DS-1-2 | 111 | AuthentiQ: Real-Time On-Device Deepfake Detection for the Internet of Multimedia Things | Raghu Iyengar (pi-labs)*; Kapardi Mallampalli (pi-labs); Hitesh Sanghvi (Directorate of Forensic Sciences) |
| 13:30-15:15 16:15-18:00 |
DS-1-3 | 148 | Real-Time Clinical Decision Support in the Trauma Room: An Edge-Native Two-Agent System for Structured Information Processing | Niklas Beck (Fraunhofer IAIS)*; Johann Jasper Schulze Buschhoff (Fraunhofer IAIS); Emmy Lai (Fraunhofer IAIS); Najmeh Mousavi (Fraunhofer IAIS); Julian Maluck (T-Systems International); Teh-Hai Julian Zheng (T-Systems International); Stefan Swoboda (T-Systems International); Benny Jörg Stein (Fraunhofer IAIS) |
| 13:30-15:15 16:15-18:00 |
DS-1-4 | 151 | Personalized Federated Fine-Tuning for Distributed Image-Text Retrieval | Yifan Zhou (shanghai university); Mengxin Yang (shanghai university); Hongmei Shi (shanghai university); Yixiao Gu (shanghai university)* |
| 13:30-15:15 16:15-18:00 |
DS-1-5 | 152 | BETA-Cap: An Interactive System for Bilingual English-Turkish Audio Captioning Best Demo Paper Candidate | Melek Tanyeri Yener (Baskent University)*; Mustafa Sert (Baskent University) |
| 13:30-15:15 16:15-18:00 |
DS-1-6 | 169 | Harness Engineering for Multi-Agent Data Visualization with Small Language Models | Xinxiao Li (Shonan Institute of Technology); Peilin Wang (Shonan Institute of Technology); Eisei Nakahara (Nihon University)* |
| 13:30-15:15 16:15-18:00 |
DS-1-7 | 178 | Designing an Intelligent Multimedia Vending Interface for Ethical-Choice Interaction | Wing Yan Jasman PANG (Technological and Higher Education Institute of Hong Kong)*; Yan Chi Wan (Technological and Higher Education Institute of Hong Kong) |
| Time | Chair | Title | Speaker | Venue |
|---|---|---|---|---|
| 11:30-12:30, Mon. 10 Aug. | Albert Ali Salah | Towards Trustworthy Multimodal Models | Prof. Mohan Kankanhalli (National University of Singapore, Singapore) |
Ballroom |
| 14:00-15:00, Mon. 10 Aug. | Supavadee Aramvith | Stable Components in High-Dimensional Data: Joint Decomposition Across Datasets | Prof. Tülay Adalı (University of Maryland, Baltimore County, USA) |
Ballroom |
| Session | Chair | Venue | Time | Note |
|---|---|---|---|---|
| RO-1 | Toshihiko Yamasaki | Ballroom | 09:00-10:30, Mon. 10 Aug. | 6 Papers |
| RO-2 | Albert Ali Salah | Ballroom | 15:30-17:15, Mon. 10 Aug. | 7 Papers |
| RP-1 | Shanshan Yao | Komes 1 | 15:30-17:45, Mon. 10 Aug. | 24 Papers |
Chair: Toshihiko Yamasaki | Venue: Ballroom | Time: 09:00-10:30, Monday, 10 August 2026
| Time | Slot | Paper ID | Paper Title | Authors |
|---|---|---|---|---|
| 09:00-09:15 | RO-1-1 | 17 | Single-View Biomechanical State Estimation for Maize Architecture Recovery | Yue Zhuo (Cornell University); Fengqi You (Cornell University)* |
| 09:15-09:30 | RO-1-2 | 18 | 3D Object Recovery for Data Tampering Attack on In-vehicle Networks | Gangtao Han (Zhengzhou University); Yurui Chen (Zhengzhou University); Song Wang (Zhengzhou University)*; Enqing Chen (Zhengzhou University); Lingling Li (Zhengzhou University of Aeronautics); Gaofeng Pan (Beijing Institute of Technology) |
| 09:30-09:45 | RO-1-3 | 45 | Boosting Vision-Language Model for Cross-Device GUI Agent Grounding and Navigation Best Student Paper Candidate | Bo-Yu Chen (National Chung Cheng University); Liang-Han Shih (National Chung Cheng University); Chen-Kuo Chiang (National Chung Cheng University)* |
| 09:45-10:00 | RO-1-4 | 47 | SpatialAgent: Spatial Question Answering with LLM Agent and Perception Models 🏆 Best Student Paper Award | Hsiang-Wei Huang (University of Washington)*; Junbin Lu (University of Washington); Jianxu Shangguan (University of Washington); Jen-Hao Cheng (University of Washington); Kuang-Ming Chen (University of Washington); Cheng-Yen Yang (University of Washington); Bahaa Alattar (University of Washington); Yi-Ru Lin (University of Washington); Pyongkun Kim (Electronics and Telecommunications Research Institute); Sangwon Kim (Electronics and Telecommunications Research Institute); Kwangju Kim (Electronics and Telecommunications Research Institute); Chung-I Huang (National Chung Hsing University); Jenq-Neng Hwang (University of Washington) |
| 10:00-10:15 | RO-1-5 | 63 | Document-Grounded Coaching Agent for Video Skill Assessment Best Paper Candidate | Masaya Fujiwaka (NEC)*; Takuro Kawada (Hosei University); Xiaotong Ji (NEC); Jianquan Liu (NEC) |
| 10:15-10:30 | RO-1-6 | 84 | Trajectory-Centered Reasoning for Counterfactual Driving Question Answering Best Paper Candidate | Xiaotong Ji (NEC Corporation)*; Masaya Fujiwaka (NEC Corporation); Jianquan Liu (NEC Corporation) |
Chair: Albert Ali Salah | Venue: Ballroom | Time: 15:30-17:15, Monday, 10 August 2026
| Time | Slot | Paper ID | Paper Title | Authors |
|---|---|---|---|---|
| 15:30-15:45 | RO-2-1 | 3 | Bridging Design and Quality: A Technical Overview of FMs in Manufacturing | Chih-Fan Hsu (Inventec Corp.)*; Yung-Huei Chiou (Inventec Corp.); Po-Han Huang (Inventec Corp.); Yu-chi Liao (Inventec Corp.); Shu Huei Yang (Inventec Corp.); Jyun-Ze Tang (Inventec Corp.); Yu-Lun Chang (Inventec Corp.); Ming-Ching Chang (University at Albany) |
| 15:45-16:00 | RO-2-2 | 24 | QuranRTT: Real-Time Domain-Specific Translation of Classical Arabic Quranic Recitation | Nagendhra Madishetti (University at Albany, State University of New York); Md Monjurul Ahsan Prodhan (University at Albany, State University of New York); Md Nour Hossain (University at Albany, State University of New York)* |
| 16:00-16:15 | RO-2-3 | 36 | Hard-Decision Bottleneck in Belief Propagation for Speaker Diarization: A Factor Graph Analysis | Snit Sanghlao (Faculty of Information and Communication Technology, Mahidol University)*; Jantarat Saenghlao (Department of Computer Engineering, KMUTT); Threrawee Sanglaow (Department of Physics, KMUTT) |
| 16:15-16:30 | RO-2-4 | 53 | Continual Learning with Support Boundary Experience Blending 🏆 Best Paper Award | Chih-Fan Hsu*; Yung-Huei Chiou (Inventec Corp.); Ming-Ching Chang (University at Albany, State University of New York); Wei-Chao Chen (Inventec Corp.) |
| 16:30-16:45 | RO-2-5 | 62 | Modeling Video Ads with Structured Semantics for CTR Prediction Best Student Paper Candidate | Yoshiaki Tanabe (The university of Tokyo)*; Shuntaro Masuda (The university of Tokyo); Gakumatsu Ryu (Septeni Japan, Inc.); Ryoichi Katsuya (Septeni Japan, Inc.); Toshihiko Yamasaki (The university of Tokyo) |
| 16:45-17:00 | RO-2-6 | 100 | VCE-DINO: Robust Video Capsule Endoscopy Anatomy Recognition via DINO Pretraining and Imbalance-Aware Learning | Chayapon Sasnarukkit (Chulalongkorn university); Natawut Nupairoj (Chulalongkorn university); Peerapon Vateekul (Chulalongkorn university)* |
| 16:45-17:15 | RO-2-7 | 40 | Bi-projection Fusion for 360 Depth Estimation Based on Coordinate Attention | Lei Li (Information Engineering University);Chenyang He (Zhengzhou University);Ying Yu (Information Engineering University);Tao Yang (Zhengzhou University);Yuhang Lu (Zhengzhou University);Yun Tie (Zhengzhou University)* |
Chair: Shanshan Yao | Venue: Komes 1 | Time: 15:30-17:45, Monday, 10 August 2026
| Time | Slot | Paper ID | Paper Title | Authors |
|---|---|---|---|---|
| 15:30-17:45 | RP-1-1 | 5 | Optimized Smart Home Energy Management Systems Based on Hybrid Method | Alexander N. Ndife (Chalmers University of Technology)*; Arinze Ijeoma (Institut für Kernphysik University of Münster); Paisarn Muneesawang (Department of Computer Engineering Mahidol University) |
| 15:30-17:45 | RP-1-2 | 16 | A Rotation-Guided and Feature-Refined Framework for Robust Nuclei Segmentation in Histopathology | Arbab Sufyan Wadood (Multimedia University)*; Mohammad Faizal Ahmad Fauzi (KPJU Healthcare University); Lai Kuan Wong (Multimedia University); Md Serajun Nabi (KPJU Healthcare University) |
| 15:30-17:45 | RP-1-3 | 20 | Hierarchical Federated Class-Incremental Learning with Client-Local Replay for Industrial IoT Intrusion Detection | Waiteng CHING (University of Aizu)*; Xiuheng LIAO (University of Aizu); Jingcheng SONG (Linyi University); Weiyu Wang (Hosei University); Zhiwei Si (University of Aizu); Chunhua SU (University of Aizu) |
| 15:30-17:45 | RP-1-4 | 25 | Multi Local Modeling: Toward Structured Symbolic Music Generation Via Local Learning | Van Hoang Bao Bach (University of Central Florida)*; Sansiri Tarnpradab (King Mongkut's University of Technology Thonburi); Kien Hua (University of Central Florida) |
| 15:30-17:45 | RP-1-5 | 28 | DeepVisionSafety: An intelligent CCTV Monitoring System for Multi-class Anomaly Recognition | Napasrapee Satittham (Department of Computer Engineering, Faculty of Engineering, Mahidol University, Nakhon Pathom, Thailand.)*; Purin Pongpanich (Department of Computer Engineering, Faculty of Engineering, Mahidol University, Nakhon Pathom, Thailand.); Kunanont Vutiprasertwong (Department of Computer Engineering, Faculty of Engineering, Mahidol University, Nakhon Pathom, Thailand.); Narit Hnoohom (Department of Computer Engineering, Faculty of Engineering, Mahidol University, Nakhon Pathom, Thailand.) |
| 15:30-17:45 | RP-1-6 | 32 | A Distributed Learning Approach for Controlled Grammar Transfer using Structure-Timbre Disentangled Audio Synthesis | Sridhar Varma Kalidindi (National Institute of Technology Andhra Pradesh); Karthick S (National Institute of Technology Tiruchirappalli)* |
| 15:30-17:45 | RP-1-7 | 33 | Attention-Enhanced EfficientNet with Anatomy-Aware Test-Time Augmentation for Robust Dental Caries Detection | Phubase Sangliamthong (Faculty of Information and Communication Technology, Mahidol University ); Aroonrat Choochue (Faculty of Information and Communication Technology, Mahidol Universit); Siriwan Suebnukarn (Faculty of Dentistry, Thammasat University ); Kanoksak Wattanachote (Mahidol University )* |
| 15:30-17:45 | RP-1-8 | 41 | Language-Conditioned Geographic Localization of Indic Speech | Raghu Iyengar (pi-labs)*; Hitesh Sanghvi (Directorate of Forensic Science ) |
| 15:30-17:45 | RP-1-9 | 46 | Detection-based Spatial Token Selection for Long Video Understanding with MLLMs | Yuta Goto (NEC Corporation)*; Kengo Ishida (NEC Corporation); Satoshi Yamazaki (NEC Corporation); Junichi Funada (NEC Corporation); Jianquan Liu (NEC Corporation) |
| 15:30-17:45 | RP-1-10 | 57 | Extreme Learning Machine Ensemble for Deep Learning-based List Decoding in Error-prone Video Transmission Systems | Anton Louise De Ocampo (Batangas State University)*; Francois-Xavier Coudoux (Institut d’Électronique de Microélectronique et de Nanotechnologie (IEMN)); Stéphane Coulombe (École de technologie supérieure (Université du Québec)); Patrick Corlay (Université Polytechnique Hauts-de-France) |
| 15:30-17:45 | RP-1-11 | 61 | Convolution and Transformer Integration for Glaucoma Classification on Enhanced Fundus Images | Chayawat Nilsumrit (Mahidol University); Narit Hnoohom (Mahidol University)* |
| 15:30-17:45 | RP-1-12 | 66 | An Empirical Study of Model-Specific Preprocessing and Focal Loss for Thai Depression Text Classification | Rawipas Samoondee (Chiang Mai University)*; Dussadee Praserttitipong (Chiang Mai University) |
| 15:30-17:45 | RP-1-13 | 70 | Real-Time Video Anomaly Detection Using YOLO Pose Estimation and CLIP-Based Semantic Scoring | Vanodhya Warnasooriya (Chulalongkorn University); Amir Hajian (Chulalongkorn University); Watchara Ruangsang (King Mongkut’s University of Technology Thonburi); Supavadee Aramvith (Chulalongkorn University)* |
| 15:30-17:45 | RP-1-14 | 71 | CARRE: Comprehensive Automated Refinement for Real Estate Using Large Language Models | Yuta Takatsuji (The University of Tokyo)*; Tatsuya Iwanari (estie, inc.); Ryosuke Katsuta (estie, inc.); Shuntaro Masuda (The University of Tokyo); Shengzhou Yi (The University of Tokyo); Toshihiko Yamasaki (The University of Tokyo) |
| 15:30-17:45 | RP-1-15 | 73 | Discriminative Correlation Perceptron-Style Neural Network for Multi-View Glaucoma Image Classification | Chakkrit Termritthikun (Naresuan University)*; Yirga Belay Muna (Naresuan University); Paisarn Muneesawang (Mahidol University); Patchara Suwanbordin (Mahidol University); Lei Gao (Wilfrid Laurier University); Kai Liu (Toronto Metropolitan University); Ling Guan (Toronto Metropolitan University) |
| 15:30-17:45 | RP-1-16 | 77 | SAM3 Adaptation for Root CT Image Segmentation | Hurriyatul Fitriyah (Adelaide University)*; Huafei Huang (University of South Australia); Ivan Lee (University of South Australia) |
| 15:30-17:45 | RP-1-17 | 82 | Beyond Global and Local: A Stacked Ensemble Approach to Symbolic Music Emotion Recognition | Hema Kanoujia (University of Central Florida)*; Hemant Kanoujia (University of Central Florida); Kien Hua (University of Central Florida) |
| 15:30-17:45 | RP-1-18 | 83 | Depth-Aware Pothole Detection Using YOLO and RT-DETR at the Edge | Md Monjurul Prodhan (University at Albany, State University of New York); Md Nour Hossain (University at Albany, State University of New York)* |
| 15:30-17:45 | RP-1-19 | 87 | Evaluating Diffusion Models for Single-Image 3D Gaussian Splatting Scene Reconstruction | Max Eskandari (York University); Xingbang Tang (York University)*; Matthew J. Kyan (York University) |
| 15:30-17:45 | RP-1-20 | 88 | A Harmless Owner-Specific Verification Scheme for DNN Copyright Protection | Jiangnan Zhu (Kyushu University)*; Yuzhang Chen (Kyushu University); Yujie Gu (Kyushu University); Minoru Kuribayashi (Tohoku University); Kouichi Sakurai (Kyushu University) |
| 15:30-17:45 | RP-1-21 | 98 | Exploring LMM-as-a-Judge for Image Harmonization Evaluation | Jeonghun Baek (The University of Tokyo)*; Eunchung Noh (Samsung Electronics) |
| 15:30-17:45 | RP-1-22 | 101 | Evaluating Symptom Fidelity in LLM-Driven Simulated Patients Using PHQ-9 and GAD-7 | Pathanin Opachalearn (Mahidol University)*; Konlakorn Wongpatikaseree (Mahidol University); Natthawut Kertkeidkachorn (Japan Advanced Institute of Science and Technology) |
| 15:30-17:45 | RP-1-23 | 105 | Characterizing Spatial Interface Metaphors for Exploratory 360° Video Retrieval in VR | Mike Knoop (Utrecht University); Wolfgang Hürst (Utrecht University)* |
| 15:30-17:45 | RP-1-24 | 107 | Improving GPT-Based Subjective Floor Plan Evaluation with Structural Cue Correction | Taro Narahara (New Jersey Institute of Technology)*; Toshihiko Yamasaki (The University of Tokyo) |
| Track | Chair | Venue | Time | Note |
|---|---|---|---|---|
| Workshop 3 | Luntian Mou and Zichun Guo | Komes 2 | 09:00-11:05, Mon. 10 Aug. | 3 Keynotes and 5 Papers |
| Workshop 4 | Komes 2 | 15:30-16:15, Mon. 10 Aug. |
Chair: Luntian Mou | Venue: Komes 2 | Time:09:00-11:05, Monday, 10 August 2026
| Time | Slot | Paper ID | Paper Title | Authors |
|---|---|---|---|---|
| 09:00-09:25 | WS-3-1 | Keynote 1 | Provenance Computation of Historical and Cultural Data | Xiaoru Yuan (Peking University) |
| 09:25-09:35 | WS-3-2 | 144 | A Value-Aligned Multimodal Agent for Culturally Grounded Textbook Illustration Generation and Evaluation | Chenrui Qin (East China University of Science and Technology); Yumiao Chen (East China University of Science and Technology)*; Yifan Yan (East China University of Science and Technology) |
| 09:35-09:45 | WS-3-3 | 145 | SensiPillow: An Embodied Multimodal Agent for Bedtime Smartphone Intervention | Qi Fang (Southern University of Science and Technology); Qianwen Tan (Southern University of Science and Technology); Kai Lin (Southern University of Science and Technology); Ziqian Bai (Southern University of Science and Technology)* |
| 09:45-10:10 | WS-3-4 | Keynote 2 | Music Large Model Shao 1.0 | Jiafeng Liu (Central Conservatory of Music) |
| 10:10-10:20 | WS-3-5 | 175 | Phygital Symbiosis and Innovation Dynamics of Location-Based Entertainment in China | Jie Pang (Peking University)*; Jiaxin Li (Peking University) |
| 10:20-10:30 | WS-3-6 | 176 | EAC2-Curation:Embodied Attention Capture and Curatorial Reconstruction | Jiaxin Li(Peking University )*; Jie Pang(Peking university); Xinyi Wang(Zhejiang University) |
| 10:30-10:55 | WS-3-7 | Keynote 3 | Jewelry and Theater: Design and Artistic Practice Spanning Physical and Digital Realities | Ting Song (Ting Museum) |
| 10:55-11:05 | WS-3-8 | 180 | From Viewing to Experiencing: An Embodied Narrative Study of Dunhuang Mural Content | Xianning Meng (Daegu University)*; Zichun Guo (Beijing University of Chemical Technology); Qiang Hao (Tsinghua University); Chang Pan (Daegu University) |
Chair: | Venue: Komes 2 | Time: 15:30-16:15, Monday, 10 August 2026
| Time | Chair | Title | Speaker | Venue |
|---|---|---|---|---|
| 10:30-11:30, Tue. 11 Aug. | Iole Moccagatta | Journey to the Next Era: AI-driven Multi-view Analysis and the Future of Spatial Intelligence | Prof. Hideo Saito (Keio University, Japan) |
Ballroom |
| Session | Chair | Venue | Time | Note |
|---|---|---|---|---|
| IP-1 | Mei-Chen Yeh | Ballroom | 09:00–10:00 / 11:30–12:15, Tue. 11 Aug. | 4/3 papers |
| IP-2 | Guan-Ming Su | Komes 1 | 09:00–10:00 / 11:30–12:15, Tue. 11 Aug. | 4/3 papers |
| IP-3 | Chih-Fan Hsu | Komes 2 | 09:00–10:00 / 11:30–12:15, Tue. 11 Aug. | 4/3 Papers |
Chair: Mei-Chen Yeh | Venue: Ballroom | Time: 09:00-10:00, Tuesday, 11 August 2026
| Time | Slot | Paper ID | Paper Title | Authors |
|---|---|---|---|---|
| 09:00-09:15 | IP-1-1 | 158 | Designing Storyboards for Immersive Video Retrieval: A Comparative Study of 360-Degree Video Visualizations | Marloes van der Haagen (Utrecht University); Wolfgang Hürst (Utrecht University)* |
| 09:15-09:30 | IP-1-2 | 109 | Self-Supervised Representation Learning Using Object-Aware Data Augmentation | Yong-Hong Lin (National Taiwan Normal University); Mei-Chen Yeh (National Taiwan Normal University)* |
| 09:30-09:45 | IP-1-3 | 173 | Force-Responsive Virtual Makeup | Ning Xu (Adeia Inc.)*; Aldis Sipolins (Adeia Inc.); Mareeta Mathai (Adeia Inc.); Mathew Adams (Adeia Inc.); Serhad Doken (Adeia Inc.) |
| 09:45-10:00 | IP-1-4 | 117 | Transformer-Based Local Path Planning for Quadruped Robots Using Multi-View Semantic Perception | Ju Wang (Virginia State University)*; Wei-Bang Chen (Virginia State University); Qian Qu (Virginia State University); Xuanchen Xiang (Virginia State University); Ahmed Mohammed (Virginia State University) |
Chair: Mei-Chen Yeh | Venue: Ballroom | Time: 11:30-12:15, Tuesday, 11 August 2026
| Time | Slot | Paper ID | Paper Title | Authors |
|---|---|---|---|---|
| 11:30-11:45 | IP-1-5 | 133 | Super-Resolution Preprocessing for Small Object Detection: An Empirical Study on VisDrone-Det at ×2 and ×4 Scales | Musa Yahaya (Chulalongkorn University)*; Amir Hajian (Chulalongkorn University); Kanapoj Ngambenjavichaikul (Chulalongkorn University); Watchara Ruangsang (King Mongkut's University of Technology Thonburi); Supavadee Aramvith (Chulalongkorn University) |
| 11:45-12:00 | IP-1-6 | 155 | A Lightweight Speaker Recognition Method with Multi-Layer Self-Distillation | Shanshan Yao (Shanxi University); Jing You (Shanxi University); Yaozhang Zhu (Shanxi Lu'an Mining (Group) Co., Ltd.); Baoning Niu (Taiyuan University of Technology)* |
| 12:00-12:15 | IP-1-7 | 149 | Physics-Guided Cascade Classification for Multi-Granularity Liquid Recognition in Dual-Energy X-ray Security Inspection | Hongping Zhi (Tsinghua University)*; Lifeng Sun (Tsinghua University); Xiaoyun Su (Suzhou ZMore Technology Co., Ltd.); Weiwei Zheng (Suzhou ZMore Technology Co., Ltd.) |
Chair: Guan-Ming Su| Venue: Komes 1 | Time: 09:00-10:00, Tuesday, 11 August 2026
| Time | Slot | Paper ID | Paper Title | Authors |
|---|---|---|---|---|
| 09:00-09:15 | IP-2-1 | 86 | Global Attribution Meets Local Representation: A Hybrid Pipeline for Binary Composer Classification | Hemant Kanoujia (University of Central Florida)*; Hema Kanoujia (University of Central Florida); Kien Hua (University of Central Florida) |
| 09:15-09:30 | IP-2-2 | 116 | Anchor Tracking and Prediction with Triplane Channel-to-Spatial Mapping for Dynamic Gaussian Splatting Compression | Birendra Kathariya (Dolby)*; Tsung-Wei Huang (Dolby Laboratories Inc.); Dae Yeol Lee (Dolby Laboratories Inc.); Fangjun Pu (Dolby Laboratories Inc.); Guan-Ming Su (Dolby Laboratories Inc.); Peng Yin (Dolby Laboratories Inc.); Gary J. Sullivan (Dolby Laboratories Inc.) |
| 09:30-09:45 | IP-2-3 | 160 | Capture-Stage Multi-Camera Placement Recommendation System for Improved 3D Scene Coverage | Dae Yeol Lee (Dolby Laboratories)*; Sandeep Bangalore Venkatesh (Dolby Laboratories); Anshul Rai (Dolby Laboratories); Tsung-Wei Huang (Dolby Laboratories); Guan-Ming Su (Dolby Laboratories) |
| 09:45-10:00 | IP-2-4 | 174 | Dynamic Range Reshaping for Gaussian Splatting | Tsung-Wei Huang (Dolby Laboratories)*; Dae Yeol Lee (Dolby Laboratories); Birendra Kathariya (Dolby Laboratories); Guan-Ming Su (Dolby Laboratories); Peng Yin (Dolby Laboratories) |
Chair: Guan-Ming Su| Venue: Komes 1 | Time: 11:30-12:15, Tuesday, 11 August 2026
| Time | Slot | Paper ID | Paper Title | Authors |
|---|---|---|---|---|
| 11:30-11:45 | IP-2-5 | 163 | Origin–Destination Proportion Analysis Using Machine Learning on Non-Overlapping Camera Videos | Sajan Kumar (Purdue University); Stanley Chien (Purdue University)*; Shail Desai (Purdue University); Yaobin Chen (Purdue University); Nathan Shellhamer (Indiana Department of Transportation) |
| 11:45-12:00 | IP-2-6 | 164 | Open-set Visual Template Discovery in Large-Scale Propaganda Image Collections | Minh Tu Nguyen (National Chung Cheng University); Yu-Yao Tseng (National Chung Cheng University); Ming-Hung Wang (National Chung Cheng University)* |
| 12:00-12:15 | IP-2-7 | 154 | Emotional Plants: When Artificial Agents Report Their Internal Affective States | Vito Vekic (Utrecht University); Albert Ali Salah (Utrecht University)*; Almila Akdag (Utrecht University) |
Chair: Chih-Fan Hsu | Venue: Komes 2 | Time: 09:00-10:00, Tuesday, 11 August 2026
| Time | Slot | Paper ID | Paper Title | Authors |
|---|---|---|---|---|
| 09:00-09:15 | IP-3-1 | 91 | Lesion-DDPM: Lesion-Enhanced 3D Diffusion for MS MRI Synthesis | Weidong Zhang (Northeastern University); Yongchan Jung (Fairleigh Dickinson University); Shafayat Mowla Anik (University of Colorado - Colorado Springs); Furen Xiao (National Taiwan University); Vasudevan Janarthanan (Fairleigh Dickinson University); Enkhzaya Chuluunbaatar (University of British Columbia); Marco Ho (British Columbia Institute of Technology); Byeong Kil Lee (University of Colorado - Colorado Springs); Jeeho Ryoo (Fairleigh Dickinson University)* |
| 09:15-09:30 | IP-3-2 | 156 | DESNet: A Multimodal Collaborative Network for Joint Driver Emotion and Stress Recognition | Luntian Mou (Beijing University of Technology)*; Wenyan Zhao (Beijing University of Technology); Shasha Mao (Xidian University); Nan Ma (Beijing University of Technology) |
| 09:30-09:45 | IP-3-3 | 167 | Edge Computing with Green Super-Resolution | Li-Heng Wang (USC)*; Yixing Wu (USC); Wei Wang (USC); C.-C. Jay Kuo (USC) |
| 09:45-10:00 | IP-3-4 | 76 | A Discriminant Multi-view Correlation Analysis Platform for Multimedia Computing | Lei Gao (Wilfrid Laurier University)*; Kai Liu (Toronto Metropolitan University); Ling Guan (Toronto Metropolitan University) |
Chair: Chih-Fan Hsu | Venue: Komes 2 | Time: 11:30-12:15, Tuesday, 11 August 2026
| Time | Slot | Paper ID | Paper Title | Authors |
|---|---|---|---|---|
| 11:30-11:45 | IP-3-5 | 161 | A Multi-Source Ultrasound Benchmark Revealing the Limits of Contemporary Self-Supervised Anomaly Detection Methods | Marco Riedenauer (University of Augsburg)*; Daniel Kienzle (University of Augsburg); Pratik Mayekar (University of Augsburg); Rainer Lienhart (University of Augsburg) |
| 11:45-12:00 | IP-3-6 | 139 | Comparison of multimodal LLM embedding models in the task of cross-modal recipe retrieval | Keisuke Gomi (The University of Electro-Communications)*; Keiji Yanai (The University of Electro-Communications) |
| 12:00-12:15 | IP-3-7 | 115 | Differentiable Proxy Learning for Adaptive Quantization Control in H.264 Video Coding | Qihan Xu (Simon Fraser University)*; Ivan Bajic (Simon Fraser University) |