Executive Summary
Cockpit localization covers language presentation, voice understanding, navigation services and device connections. In addition to module-level checks, evaluation needs to establish whether users can complete the intended operation under specified conditions. For a multi-market project, the applicability of each result depends on participant characteristics, vehicle build and local service configuration.
This report draws on published owner surveys, platform documentation and writing-system references to examine localization testing for Indonesia, Thailand, Vietnam, Malaysia and Singapore. It uses individual tasks as the unit of evaluation, recording completion, stage-specific errors and recovery for diagnosis and acceptance review.
The proposed test arrangements require validation on the intended configuration. The owner-survey evidence comes from Thailand and does not represent all five markets. No target-vehicle testing was conducted for this report, and no vehicle scores, market rankings or benefit estimates are provided.
1. Evaluation scope
A localization test record should identify the market and location, user profile, vehicle build, phone and service configuration, operating state and task. Together, these fields define where the result applies. Language proficiency should be established during recruitment; nationality alone does not describe proficiency or operating experience.
Testing can begin with configurations covered by the product's support commitments, adding variations according to task risk. Exhaustive testing of every combination is not assumed. Reuse of a result in another market requires a comparison of language materials, services and equipment.
Platform differences also need separate records. Android Auto provides a vehicle experience through a connected phone, while Android Automotive OS is built into the vehicle. Similar tasks therefore have different application locations and dependencies. Results should be retained separately for each platform. Android platform definitions
Table 1 lists the proposed scope fields. Unconfirmed configurations remain open items.
Proposed scope template | Not market-research findings
| Dimension | Record | Evaluation purpose |
|---|---|---|
| Market/location | Launch country and target cities or usage locations | Bound place-name, route and service-coverage checks |
| User profile | Language/proficiency, first or experienced use, driver/passenger role | Recruit and segment by capability, not nationality alone |
| Vehicle build | Model, hardware/software, display and control layout | Bind results to a build; read steering position from actual configuration |
| Phone/services | Phone/OS, connection mode, maps and account environment | Separate native and phone-connected services; identify dependencies |
| Operating state | Parked/driving, network/acoustics, concurrent calls or media | Make conditions reproducible; use professional driving-test procedures |
| Target task | Initial state, goal, success and stopping conditions | Agree outcomes before testing instead of counting features |
The scope record informs local language review, recruitment, equipment selection and test scripts. Cross-market differences can then be examined as specific language, location, service or configuration differences.
Without interviews or usage records, phone connection, destination search and climate-control access remain candidate tasks. The available material does not establish their demand ranking in each market. Priorities require further target-user evidence, product scoping and task-risk assessment.
2. Task decomposition and diagnosis
A hypothetical voice-navigation task is used to illustrate the method: a user names a local destination, the system retrieves a place, and navigation starts after confirmation. This scenario does not describe an observed brand incident or assume that users prefer voice input.
The operation can be divided into activation, understanding, resolution, confirmation and execution. Automatic speech recognition (ASR) handles transcription, and natural language understanding (NLU) handles intent and entities. Subsequent stages involve point-of-interest (POI) retrieval, candidate confirmation and vehicle integration. Table 2 assigns outcomes, diagnostic measures and module responsibilities to these stages.
Proposed method | Not observed results or verified capabilities
| Stage | Required outcome | Diagnostic measure | Owner |
|---|---|---|---|
| 01 Activate | Enter the intended interaction state | Activation success; false activations per test hour | Wake/input team |
| 02 Understand | Resolve intent and entities, not just a transcript | Intent/entity accuracy; transcription error as diagnosis | ASR/NLU |
| 03 Resolve | Retrieve a destination that fulfils the user's intent | Correct-candidate availability and resolved selection | Maps/search |
| 04 Confirm | Let the user understand, correct or cancel | Unassisted first confirmation; repair turns | HMI/dialogue |
| 05 Execute | Start the intended navigation and report its state | End-to-end completion, latency and recovery | Vehicle integration/navigation |
When transcription is correct but the place is wrong, retrieval scope, candidate ranking and confirmation information warrant examination. When the place is correct but navigation does not start, execution requests and state feedback need inspection. These are possible explanations until test records establish the cause.
One task owner can consolidate the complete result, with module teams supplying inputs, outputs and error records. Module and whole-task findings should be cross-referenced so that failures at handover points remain visible.
3. Published evidence and implications for testing
Owner surveys describe reported problems; platform documentation defines systems and design requirements; task tests on a specified configuration evaluate the proposed delivery. The sources below inform test selection. They do not establish that the intended vehicle has the same defects.
J.D. Power's 2025 Thailand IQS Volume 1 surveyed 4,721 new-vehicle owners between December 2024 and February 2025. The release identifies infotainment as the most problematic category and lists Bluetooth connectivity, device power/charging and touchscreen response among the reported issues. J.D. Power release
These findings support examining device connections and interaction procedures. The study does not attribute the issues to inadequate localization, and this report does not use them to assess a particular brand or another country. The following test areas are considered in relation to the navigation task.
3.1 Language presentation and input
Linguistic quality assurance (LQA) covers terminology, meaning and consistency. Device testing covers rendering, input and correction. Reviewed translations still need inspection on the intended display.
W3C's Thai layout draft describes phrase-based spacing and vowel marks above or below consonants. W3C Thai draft These features suggest checking line breaks and clipped marks. The document does not establish a defect in any particular cockpit.
Navigation checks can cover complete destination names, information distinguishing candidates, consistency between spoken and visual output, and deletion, editing and cancellation. Records should retain string identifiers, locale, font and software version, captures and local-review findings. Day/night themes, display configuration and reachability require inspection on the intended vehicle.
3.2 Response time and driver attention
Google's driving-interaction guidance specifies interface responses within 0.25 seconds of input and a processing indication when content loading exceeds 2 seconds. It also calls for interruptible, resumable interactions and priority for driving information. Google interaction guidance
These timing requirements concern the relevant interface feedback. They are not cloud-task completion deadlines or Southeast Asian legal limits. Navigation testing should record input-to-first-perceptible-feedback and input-to-navigation-start separately.
Further checks include parked/driving restrictions, candidate-list comprehension and priorities among navigation, vehicle warnings, calls and media. Hands-free voice interaction still imposes cognitive demands. Extended dialogue, sign-in and long lists require human-factors review; driving-related trials need qualified personnel and appropriate controlled conditions.
3.3 Cabin acoustics
J.D. Power's 2025 Thailand IQS Volume 2 reports four noise categories. Fieldwork covered June–October 2025 and included 4,832 new-vehicle owners. J.D. Power release
Figure 1 uses PP100, or reported problems per 100 vehicles. Road noise has the highest value among these four disclosed categories. PP100 is neither an affected-owner percentage nor a sound-pressure measurement, and the comparison does not constitute a complete cockpit-UX ranking.
June–October 2025 · n=4,832 · Problems per 100 vehicles (PP100) · PP100 is problems per 100 vehicles, not the share of affected owners.
www.jdpower.comThe figures describe owner-reported noise problems. The release provides no recognition-performance results, so the relationship between noise and recognition failure, and the benefits of improving voice performance, cannot be estimated from this study. Summing the values would not produce an affected-owner share.
For configurations using in-vehicle voice interaction, road noise, HVAC airflow, passenger speech and media playback can be specified as controlled test conditions. Transcription, entity resolution and recovery can then be compared for the same task. Any effect of those conditions requires confirmation through comparative testing.
The Thailand data serves only as a reference for test conditions here. It does not validate the overall evaluation method or form part of a regional trend assembled from differently scoped earlier studies.
3.4 Connections, service states and recovery
Connection and service tests can cover initial pairing, automatic reconnection, device switching, incoming calls, network loss, expired sign-in and service timeout. A destination remaining on screen does not establish data freshness; completed pairing also requires a check of the actual audio route.
Android's offline-first guidance requires applications using that architecture to support offline reads at minimum. Android offline-first guidance The requirement has a specific architectural context. Which cockpit functions remain available offline, and how they degrade, depends on product scope and verified capabilities.
For navigation, checks can include cache-age labels, destination retention after failure, duplicate execution on recovery and renewed confirmation. Error feedback should describe the current state and actions still available to the user.
Voice, location and account data also require documented purposes, processing locations, retention and deletion arrangements. These need review by the relevant market-compliance owner. General architecture guidance is insufficient to establish cross-border data compliance.
4. Measures and statistical definitions
Task completion describes the outcome. Stage-specific errors, response time and recovery describe the process. PP100, completion rates and perceived difficulty have different definitions and denominators; they are reported separately here, without a composite UX score.
A consistent task trial is proposed: one participant starts a scripted task from its specified initial state. Repetition, reselection or reconnection within that trial is a retry and does not add a denominator entry. Assistance, timeout and abandonment are recorded under predefined rules.
Invalid trials, such as test-equipment failures, may be excluded under rules agreed beforehand, with counts and reasons disclosed. A user's failure to complete a task is a result, not invalid data. Table 3 provides proposed measurement definitions. Acceptance thresholds require agreement before formal testing, with reference to task risk, applicable requirements, baseline research and product commitments.
Proposed method | Not observed results or verified capabilities
| Metric | Operational definition | Reporting requirement |
|---|---|---|
| Task completion | Unassisted completed trials before predefined stopping conditions / all valid task trials | Report successes, trials and participants; within-trial retries do not add denominator entries |
| Retry-free completion | Valid trials completed without repetition, reselection or reconnection / all valid task trials | Use the same denominator as eventual completion; retries are events within a trial |
| Latency P50/P95 | Measure input-to-feedback and input-to-execution separately; P95 is the 95th percentile of the relevant latency sample | Declare timestamps, sample size and method; report failures/timeouts separately |
| Recovery success | Fault-injection trials restored to the predefined usable state within agreed time/assistance rules / all valid fault-injection trials | Split connectivity loss, expired sign-in and device disconnection |
| Visual attention | Off-road glance duration, frequency and total exposure alongside task outcome | Protocol owned by human-factors/safety specialists; no generic UX score as a safety substitute |
| Perceived task difficulty | Consistent post-task wording and response scale | Retain distributions, scale and feedback from unsuccessful participants |
| Release-blocking defects | Severity defined in advance by harm, task blockage and recoverability | Review separately; do not average away critical issues with satisfaction scores |
Results should include participant counts, valid trials and successes, retaining task, market-configuration and language-proficiency groups. Analysis needs to preserve repeated measurements from the same participant. First-use and experienced-use results are reported separately.
Latency statistics require event definitions, sample sizes and percentile methods. P50 is the median and P95 the 95th percentile. Failures and timeouts should be disclosed alongside successful-trial latency. When observations are insufficient, individual values and the instability of tail percentiles should be reported.
Perceived difficulty records participants' assessments and can help explain operating problems. Safety assessment remains separate. Critical defects need individual disposition and cannot be offset by a high average satisfaction score.
5. Implementation and acceptance review
A project can begin with scoping and baseline testing in one selected market before determining test arrangements elsewhere. Timing depends on vehicle maturity, local materials and participant recruitment. The available information does not support a universal schedule.
During scoping, product and local-research owners agree tasks and supported configurations and resolve missing user and service information. Language review, target-user tests and device-record analysis follow. Exploratory research identifies issues; its sample results do not directly establish market incidence.
Issue records should contain expected and actual outcomes, conditions, build details and evidence. Unverified causes remain hypotheses until investigation establishes corrective ownership. Identified safety or compliance risks and core-task blockers are recommended for earlier attention. Efficiency, recovery and presentation issues can be prioritized using recurrence and impact evidence. Unknown frequency remains marked as unverified.
After correction, failed tasks should be retested under comparable scripts, configurations and sample conditions, including related procedures. Records should identify build changes and retain regression evidence; a single successful demonstration is insufficient for acceptance.
The acceptance review lists passed, failed and untested combinations separately, with safety and compliance matters referred to the relevant specialists. Each conclusion should trace to user characteristics, configuration, script, inputs/outputs and a responsible owner. These records also identify checks that need repeating after a language addition, map update or software upgrade.
6. Applicability and limitations
This report is a methodological analysis of localization testing and acceptance. Its published owner evidence comes principally from Thailand, and the platform documents each have a defined scope. They inform test selection but do not establish preferences across all five markets or confirm a vehicle's capabilities.
Application to a project requires a launch market, target-user definition, vehicle build, priority tasks and baseline data. No vehicle trials were conducted for this report. It supplies no vehicle scores, project cases or benefit estimates, and the dated surveys do not represent the latest regional picture in 2026.
Platform guidance and the W3C draft are technical references. Local regulatory requirements and vehicle safety certification require separate confirmation.