Smart City Governance VLM Benchmark: Kaohsiung Sovereign AI
- Linker Vision

- 1 day ago
- 4 min read
As Visual Language Models (VLMs) rapidly evolve in the global artificial intelligence landscape, their ability to integrate image analysis with semantic reasoning is becoming increasingly prominent.
However, existing major international VLM evaluation benchmarks are primarily designed around controlled scenarios required for academic research, making it difficult for them to directly reflect the diverse, real-time, and highly uncertain challenges encountered in practical urban governance. This disparity means that existing benchmarks often fail to adequately measure the effectiveness and reliability of AI models in real-world environments, making them less suitable as ability validation tools for local governments deploying AI systems.
To bridge the gap between traditional benchmarks and the demands of urban governance, the Kaohsiung City Government and Linker Vision are jointly releasing "Smart City Governance VLM Benchmark: Kaohsiung Sovereign AI" Version 2.0. This report proposes a VLM assessment standard specifically designed for smart city application scenarios, aiming to serve as a crucial validation basis before AI systems are deployed in the urban governance domain.
Rooted in Local Practice, Covering Diverse and Complex Scenarios
The core principle behind the establishment of this benchmark is to be rooted in local application logic. All question types, data sources, and task scenarios are derived from the actual governance needs of Kaohsiung City. The data incorporates diverse governance cases and visual materials provided by five municipal departments and three state-owned enterprises. These scenarios are systematically categorized into ten specific, clear, and non-overlapping urban governance domains, and expanded into 100+ L1 scenarios, 536 L2 test questions, and 4,475 effective test samples — 85.3% of them classified as high data-acquisition difficulty.

Unlike most international benchmarks that focus predominantly on "static scenarios", this benchmark emphasizes examining the model's multi-level semantic comprehension and its ability to infer event development context within "dynamic scenarios". Furthermore, the scenario design incorporates different time periods (daytime 75%, twilight 2%, nighttime 23%), weather conditions benchmarked against ten years of local meteorological statistics (clear 52%, overcast 20%, rain 28%), and diverse spatial environments to comprehensively assess the VLM's stability and adaptability across different governance scenes.
Beyond Recognition: Assessing the Model’s Ability to Understand "Governance Logic"
A key innovation of this assessment standard is its multi-level question design (L1/L2), which closely mirrors the urban decision-making logic. L1 defines the governance scenario and its semantic context, while L2 carries the questions that are scored — moving from determining "if an event has occurred" to inferring the "event type," "scope of impact," and "if response measures need to be activated". This design reinforces the completeness of the model's inference chain, ensuring its operational logic aligns with the city’s actual response and decision-making processes.
The benchmark covers four main categories of question types, corresponding to different inference and judgment requirements: Binary classification (e.g., Yes/No), Categorical classification (e.g., types of accidents), Sequential/Ordinal questions (used for ranking severity or levels), and Numerical questions (for data recognition/estimation). For sequential/ordinal and numerical questions, the scoring system employs a flexible penalty mechanism based on the error ratio, ensuring the assessment results accurately reflect subtle differences in the model's ability to handle tasks requiring sequence or numerical estimation. For rare, high-sensitivity events such as fire, the public test set deliberately includes only no-event imagery, in order to measure false-positive control — whether a model avoids raising alerts when nothing has happened.

General-Purpose Models vs. a Locally Fine-Tuned Model
Two general-purpose pretrained models and the project's own task-oriented model were evaluated under identical test data, prompts, and scoring conditions:
Governance Domain (Avg. Recall) | VILA 1.5 40B | LLaVA OneVision 7B | Kaohsiung Smart City VLM |
Traffic Flow Management | 0.437 | 0.521 | 0.864 |
Weather and Environmental Management | 0.716 | 0.936 | 0.979 |
Event and Activity Management | 0.454 | 0.656 | 0.925 |
Venues and Public Buildings | 0.519 | 0.629 | 0.957 |
Water Conservancy and Drainage Facilities | 0.554 | 0.805 | 0.935 |
Environmental Protection and Pollution Monitoring | 0.669 | 0.731 | 0.940 |
Public Transportation Facilities | 0.491 | 0.647 | 0.902 |
Energy and Power Facilities | 0.365 | 0.577 | 0.949 |
Structural Engineering and Construction Safety | 0.426 | 0.832 | 0.950 |
Factory and Port Area Facilities | 0.581 | 0.617 | 0.893 |
Total Average | 0.521 | 0.696 | 0.929 |
Model capability scores across ten urban governance domains, demonstrating the benchmark’s high discriminative power in complex scenarios.
The pattern is the finding. General-purpose models hold up where visual cues are explicit — LLaVA OneVision 7B reaches 0.936 in Weather and Environmental Management — but drop sharply in domains that depend on local governance semantics, such as Traffic Flow Management (0.521) and Energy and Power Facilities (0.577). The locally fine-tuned model is far more consistent across all ten. This is the case for sovereign AI in practice: local infrastructure patterns, administrative rules, and public-safety criteria are not reliably learnable from internet-scale data. These figures reflect baseline performance under the specified conditions and are not equivalent to post-deployment results.
Regarding data quality, all image data undergoes a systematic three-stage labeling quality assurance process—initial labeling, quality assurance review, and supervisory spot check or full check—to ensure the dataset possesses high accuracy and consistency. Furthermore, images involving public spaces adhere to privacy protection processing principles, such as blurring or masking faces and license plates, to comply with regulations and ethics.
Continuous Evolution: Building a Reference for Cross-City Collaboration
The "Smart City Governance VLM Benchmark" is not merely a dataset but a complete AI evaluation framework, whose core value lies in being grounded in genuine governance needs. Its results demonstrate that the framework can reveal meaningful differences in how models interpret complex governance scenarios, on a structured and repeatable basis.
In the future, the benchmark will continue its maintenance and localization updates, including incorporating more data and complex scenarios. Following public release, it is positioned as a practical reference for other cities and governance organizations examining multimodal-model adoption, with engagement proceeding case by case, and as a case study for international research exchange — providing a basis for exploring alignment with established evaluation frameworks over time. The goal is a capability validation tool that supports local governance first, and becomes exportable internationally as that evidence base matures.
If you wish to delve deeper into the complete content, detailed assessment methodology, and data results of the "Smart City Governance VLM Benchmark: Kaohsiung Sovereign AI" report, please click here to read the full guidance document: https://data.kcg.gov.tw/resource/069d4d72-a893-449f-998c-fd1c2d5c08f5



Comments