research brief

Cockpit localization in Southeast Asia: evaluation scope, testing and acceptance

A methodological review of cockpit localization evaluation for Southeast Asian markets, covering test scope, task diagnosis, measurement and acceptance. Published owner-survey evidence comes from Thailand; no target-vehicle testing was conducted.

AI concept illustration: an unbranded cockpit with blue navigation, viewed from the rear seats toward a rain-wet tropical street. Not a real vehicle test.

Executive Summary

Cockpit localization covers language presentation, voice understanding, navigation services and device connections. In addition to module-level checks, evaluation needs to establish whether users can complete the intended operation under specified conditions. For a multi-market project, the applicability of each result depends on participant characteristics, vehicle build and local service configuration.

This report draws on published owner surveys, platform documentation and writing-system references to examine localization testing for Indonesia, Thailand, Vietnam, Malaysia and Singapore. It uses individual tasks as the unit of evaluation, recording completion, stage-specific errors and recovery for diagnosis and acceptance review.

The proposed test arrangements require validation on the intended configuration. The owner-survey evidence comes from Thailand and does not represent all five markets. No target-vehicle testing was conducted for this report, and no vehicle scores, market rankings or benefit estimates are provided.

1. Evaluation scope

A localization test record should identify the market and location, user profile, vehicle build, phone and service configuration, operating state and task. Together, these fields define where the result applies. Language proficiency should be established during recruitment; nationality alone does not describe proficiency or operating experience.

Testing can begin with configurations covered by the product's support commitments, adding variations according to task risk. Exhaustive testing of every combination is not assumed. Reuse of a result in another market requires a comparison of language materials, services and equipment.

Platform differences also need separate records. Android Auto provides a vehicle experience through a connected phone, while Android Automotive OS is built into the vehicle. Similar tasks therefore have different application locations and dependencies. Results should be retained separately for each platform. Android platform definitions

Table 1 lists the proposed scope fields. Unconfirmed configurations remain open items.

Table 1 | Test scope

Proposed scope template | Not market-research findings

DimensionRecordEvaluation purpose
Market/locationLaunch country and target cities or usage locationsBound place-name, route and service-coverage checks
User profileLanguage/proficiency, first or experienced use, driver/passenger roleRecruit and segment by capability, not nationality alone
Vehicle buildModel, hardware/software, display and control layoutBind results to a build; read steering position from actual configuration
Phone/servicesPhone/OS, connection mode, maps and account environmentSeparate native and phone-connected services; identify dependencies
Operating stateParked/driving, network/acoustics, concurrent calls or mediaMake conditions reproducible; use professional driving-test procedures
Target taskInitial state, goal, success and stopping conditionsAgree outcomes before testing instead of counting features

The scope record informs local language review, recruitment, equipment selection and test scripts. Cross-market differences can then be examined as specific language, location, service or configuration differences.

Without interviews or usage records, phone connection, destination search and climate-control access remain candidate tasks. The available material does not establish their demand ranking in each market. Priorities require further target-user evidence, product scoping and task-risk assessment.

2. Task decomposition and diagnosis

A hypothetical voice-navigation task is used to illustrate the method: a user names a local destination, the system retrieves a place, and navigation starts after confirmation. This scenario does not describe an observed brand incident or assume that users prefer voice input.

The operation can be divided into activation, understanding, resolution, confirmation and execution. Automatic speech recognition (ASR) handles transcription, and natural language understanding (NLU) handles intent and entities. Subsequent stages involve point-of-interest (POI) retrieval, candidate confirmation and vehicle integration. Table 2 assigns outcomes, diagnostic measures and module responsibilities to these stages.

Table 2 | Voice-navigation stages and ownership

Proposed method | Not observed results or verified capabilities

StageRequired outcomeDiagnostic measureOwner
01 ActivateEnter the intended interaction stateActivation success; false activations per test hourWake/input team
02 UnderstandResolve intent and entities, not just a transcriptIntent/entity accuracy; transcription error as diagnosisASR/NLU
03 ResolveRetrieve a destination that fulfils the user's intentCorrect-candidate availability and resolved selectionMaps/search
04 ConfirmLet the user understand, correct or cancelUnassisted first confirmation; repair turnsHMI/dialogue
05 ExecuteStart the intended navigation and report its stateEnd-to-end completion, latency and recoveryVehicle integration/navigation

When transcription is correct but the place is wrong, retrieval scope, candidate ranking and confirmation information warrant examination. When the place is correct but navigation does not start, execution requests and state feedback need inspection. These are possible explanations until test records establish the cause.

One task owner can consolidate the complete result, with module teams supplying inputs, outputs and error records. Module and whole-task findings should be cross-referenced so that failures at handover points remain visible.

3. Published evidence and implications for testing

Owner surveys describe reported problems; platform documentation defines systems and design requirements; task tests on a specified configuration evaluate the proposed delivery. The sources below inform test selection. They do not establish that the intended vehicle has the same defects.

J.D. Power's 2025 Thailand IQS Volume 1 surveyed 4,721 new-vehicle owners between December 2024 and February 2025. The release identifies infotainment as the most problematic category and lists Bluetooth connectivity, device power/charging and touchscreen response among the reported issues. J.D. Power release

These findings support examining device connections and interaction procedures. The study does not attribute the issues to inadequate localization, and this report does not use them to assess a particular brand or another country. The following test areas are considered in relation to the navigation task.

3.1 Language presentation and input

Linguistic quality assurance (LQA) covers terminology, meaning and consistency. Device testing covers rendering, input and correction. Reviewed translations still need inspection on the intended display.

W3C's Thai layout draft describes phrase-based spacing and vowel marks above or below consonants. W3C Thai draft These features suggest checking line breaks and clipped marks. The document does not establish a defect in any particular cockpit.

Navigation checks can cover complete destination names, information distinguishing candidates, consistency between spoken and visual output, and deletion, editing and cancellation. Records should retain string identifiers, locale, font and software version, captures and local-review findings. Day/night themes, display configuration and reachability require inspection on the intended vehicle.

3.2 Response time and driver attention

Google's driving-interaction guidance specifies interface responses within 0.25 seconds of input and a processing indication when content loading exceeds 2 seconds. It also calls for interruptible, resumable interactions and priority for driving information. Google interaction guidance

These timing requirements concern the relevant interface feedback. They are not cloud-task completion deadlines or Southeast Asian legal limits. Navigation testing should record input-to-first-perceptible-feedback and input-to-navigation-start separately.

Further checks include parked/driving restrictions, candidate-list comprehension and priorities among navigation, vehicle warnings, calls and media. Hands-free voice interaction still imposes cognitive demands. Extended dialogue, sign-in and long lists require human-factors review; driving-related trials need qualified personnel and appropriate controlled conditions.

3.3 Cabin acoustics

J.D. Power's 2025 Thailand IQS Volume 2 reports four noise categories. Fieldwork covered June–October 2025 and included 4,832 new-vehicle owners. J.D. Power release

Figure 1 uses PP100, or reported problems per 100 vehicles. Road noise has the highest value among these four disclosed categories. PP100 is neither an affected-owner percentage nor a sound-pressure measurement, and the comparison does not constitute a complete cockpit-UX ranking.

Figure 1 | Four reported noise issues in Thailand

June–October 2025 · n=4,832 · Problems per 100 vehicles (PP100) · PP100 is problems per 100 vehicles, not the share of affected owners.

  • Road noise11 PP100
  • Wind noise6 PP100
  • Suspension noise3 PP100
  • Window operation noise2 PP100
www.jdpower.com

The figures describe owner-reported noise problems. The release provides no recognition-performance results, so the relationship between noise and recognition failure, and the benefits of improving voice performance, cannot be estimated from this study. Summing the values would not produce an affected-owner share.

For configurations using in-vehicle voice interaction, road noise, HVAC airflow, passenger speech and media playback can be specified as controlled test conditions. Transcription, entity resolution and recovery can then be compared for the same task. Any effect of those conditions requires confirmation through comparative testing.

The Thailand data serves only as a reference for test conditions here. It does not validate the overall evaluation method or form part of a regional trend assembled from differently scoped earlier studies.

3.4 Connections, service states and recovery

Connection and service tests can cover initial pairing, automatic reconnection, device switching, incoming calls, network loss, expired sign-in and service timeout. A destination remaining on screen does not establish data freshness; completed pairing also requires a check of the actual audio route.

Android's offline-first guidance requires applications using that architecture to support offline reads at minimum. Android offline-first guidance The requirement has a specific architectural context. Which cockpit functions remain available offline, and how they degrade, depends on product scope and verified capabilities.

For navigation, checks can include cache-age labels, destination retention after failure, duplicate execution on recovery and renewed confirmation. Error feedback should describe the current state and actions still available to the user.

Voice, location and account data also require documented purposes, processing locations, retention and deletion arrangements. These need review by the relevant market-compliance owner. General architecture guidance is insufficient to establish cross-border data compliance.

4. Measures and statistical definitions

Task completion describes the outcome. Stage-specific errors, response time and recovery describe the process. PP100, completion rates and perceived difficulty have different definitions and denominators; they are reported separately here, without a composite UX score.

A consistent task trial is proposed: one participant starts a scripted task from its specified initial state. Repetition, reselection or reconnection within that trial is a retry and does not add a denominator entry. Assistance, timeout and abandonment are recorded under predefined rules.

Invalid trials, such as test-equipment failures, may be excluded under rules agreed beforehand, with counts and reasons disclosed. A user's failure to complete a task is a result, not invalid data. Table 3 provides proposed measurement definitions. Acceptance thresholds require agreement before formal testing, with reference to task risk, applicable requirements, baseline research and product commitments.

Table 3 | Cockpit UX measurement dictionary

Proposed method | Not observed results or verified capabilities

MetricOperational definitionReporting requirement
Task completionUnassisted completed trials before predefined stopping conditions / all valid task trialsReport successes, trials and participants; within-trial retries do not add denominator entries
Retry-free completionValid trials completed without repetition, reselection or reconnection / all valid task trialsUse the same denominator as eventual completion; retries are events within a trial
Latency P50/P95Measure input-to-feedback and input-to-execution separately; P95 is the 95th percentile of the relevant latency sampleDeclare timestamps, sample size and method; report failures/timeouts separately
Recovery successFault-injection trials restored to the predefined usable state within agreed time/assistance rules / all valid fault-injection trialsSplit connectivity loss, expired sign-in and device disconnection
Visual attentionOff-road glance duration, frequency and total exposure alongside task outcomeProtocol owned by human-factors/safety specialists; no generic UX score as a safety substitute
Perceived task difficultyConsistent post-task wording and response scaleRetain distributions, scale and feedback from unsuccessful participants
Release-blocking defectsSeverity defined in advance by harm, task blockage and recoverabilityReview separately; do not average away critical issues with satisfaction scores

Results should include participant counts, valid trials and successes, retaining task, market-configuration and language-proficiency groups. Analysis needs to preserve repeated measurements from the same participant. First-use and experienced-use results are reported separately.

Latency statistics require event definitions, sample sizes and percentile methods. P50 is the median and P95 the 95th percentile. Failures and timeouts should be disclosed alongside successful-trial latency. When observations are insufficient, individual values and the instability of tail percentiles should be reported.

Perceived difficulty records participants' assessments and can help explain operating problems. Safety assessment remains separate. Critical defects need individual disposition and cannot be offset by a high average satisfaction score.

5. Implementation and acceptance review

A project can begin with scoping and baseline testing in one selected market before determining test arrangements elsewhere. Timing depends on vehicle maturity, local materials and participant recruitment. The available information does not support a universal schedule.

During scoping, product and local-research owners agree tasks and supported configurations and resolve missing user and service information. Language review, target-user tests and device-record analysis follow. Exploratory research identifies issues; its sample results do not directly establish market incidence.

Issue records should contain expected and actual outcomes, conditions, build details and evidence. Unverified causes remain hypotheses until investigation establishes corrective ownership. Identified safety or compliance risks and core-task blockers are recommended for earlier attention. Efficiency, recovery and presentation issues can be prioritized using recurrence and impact evidence. Unknown frequency remains marked as unverified.

After correction, failed tasks should be retested under comparable scripts, configurations and sample conditions, including related procedures. Records should identify build changes and retain regression evidence; a single successful demonstration is insufficient for acceptance.

The acceptance review lists passed, failed and untested combinations separately, with safety and compliance matters referred to the relevant specialists. Each conclusion should trace to user characteristics, configuration, script, inputs/outputs and a responsible owner. These records also identify checks that need repeating after a language addition, map update or software upgrade.

6. Applicability and limitations

This report is a methodological analysis of localization testing and acceptance. Its published owner evidence comes principally from Thailand, and the platform documents each have a defined scope. They inform test selection but do not establish preferences across all five markets or confirm a vehicle's capabilities.

Application to a project requires a launch market, target-user definition, vehicle build, priority tasks and baseline data. No vehicle trials were conducted for this report. It supplies no vehicle scores, project cases or benefit estimates, and the dated surveys do not represent the latest regional picture in 2026.

Platform guidance and the W3C draft are technical references. Local regulatory requirements and vehicle safety certification require separate confirmation.

Source evidence

Research | SEA Cockpit Lab