Screening Interface Usability With Multi-Modal LLMs: A Case Study on BOTMA Dashboards

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Usability testing of engineering dashboards typically depends on time-consuming expert walkthroughs or formal questionnaires such as the System Usability Scale (SUS). We outline an exploratory study that treats a multi-modal large language model as an automated usability assessor for a browser-based Bearing-Only Target-Motion-Analysis (BOTMA) simulation interface. A curated set of dashboard screenshots will be scored both by human and by multi-modal LLMs through a structured prompt that asks for a comparable rating and a brief justification. By examining statistical agreement and qualitative overlap between human and model scores, we aim to gauge whether LLM output is sufficiently consistent to replace - or at least screen - early-stage human evaluations. Because the calibration in this study is learned in-sample on the same set of screenshots used for evaluation, the results should be interpreted as a feasibility demonstration of screening rather than evidence of a fully generalizable replacement for expert usability assessment. Although the approach relies on static imagery rather than interactive sessions, it promises a lightweight, repeatable check that can be inserted into agile development cycles for complex simulation software. The study is expected to provide initial evidence on the practicality, limitations, and required prompt design for VLM-based usability auditing in technical web applications. © 2026 The Authors.

키워드

BOTMAdashboard evaluationLikert calibrationmulti-modal large language modelsusability testingvision-language models
제목
Screening Interface Usability With Multi-Modal LLMs: A Case Study on BOTMA Dashboards
저자
Lee, JoonwooKim, SeunghoLee, Scott Uk-Jin
DOI
10.1109/ACCESS.2026.3672166
발행일
2026-03
유형
Article
저널명
IEEE Access
14
페이지
39451 ~ 39460