e-ISSN 3135-3169
iFuture is a premier open-access journal published by Tsinghua University Press on the SciOpen platform, with academic support from the Institute for Interdisciplinary Information Sciences at Tsinghua University. Led by Turing Award Laureate Prof. Andrew Chi-Chih Yao as Editor-in-Chief, the journal is the core component of the AI Open Alliance. Its core mission is to break through AI’s theoretical bottlenecks and foundational infrastructure.
Jacob Arndt、Abhishek Potnis、Alexandre Sorokine
Electronic Navigational Charts (ENCs) are geospatial vector datasets used in maritime navigation systems that represent hydrographic and navigational information such as depths, navigational aids, traffic schemes, and hazards. A major challenge for hydrographic offices is determining whether a given chart change poses a critical or non-critical risk to maritime safety. Existing workflows rely heavily on manual review and verification, which is labor-intensive, scales poorly with the volume of incoming chart updates, and introduces inter-analyst inconsistencies. To address this challenge, we propose a method for automated classification of ENC changes. We establish a baseline encoding scheme to translate complex vector data changes into a structured tabular format for classification models. The two crucial components of the encoding scheme include a spatial context encoder to enrich the change representations with surrounding geographic features, and an ENC attribute encoder to represent nuanced attribute-value descriptions of the modified objects. We evaluate the proposed approach across two distinct operational datasets, comprising 1,308 chart pairs containing over 100,000 individual chart modifications. Tuned gradient-boosted trees leveraging the proposed encoding schemes achieve accuracies of 90% and 94% on the two datasets, yielding a 5-7% improvement over default hyperparameterized models trained on encodings without spatial context and attribute embeddings. These results demonstrate the viability of integrating machine learning into operational geospatial pipelines to improve ENC maintenance and enhance maritime safety. Finally, our experiments demonstrate the effectiveness of simple location and spatial aggregation methods, providing a foundation for evaluating more sophisticated spatial representation learning techniques for this application.
2026-09-12
e-ISSN 3135-3169
iFuture is a premier open-access journal published by Tsinghua University Press on the SciOpen platform, with academic support from the Institute for Interdisciplinary Information Sciences at Tsinghua University. Led by Turing Award Laureate Prof. Andrew Chi-Chih Yao as Editor-in-Chief, the journal is the core component of the AI Open Alliance. Its core mission is to break through AI’s theoretical bottlenecks and foundational infrastructure.
Jacob Arndt、Abhishek Potnis、Alexandre Sorokine
Electronic Navigational Charts (ENCs) are geospatial vector datasets used in maritime navigation systems that represent hydrographic and navigational information such as depths, navigational aids, traffic schemes, and hazards. A major challenge for hydrographic offices is determining whether a given chart change poses a critical or non-critical risk to maritime safety. Existing workflows rely heavily on manual review and verification, which is labor-intensive, scales poorly with the volume of incoming chart updates, and introduces inter-analyst inconsistencies. To address this challenge, we propose a method for automated classification of ENC changes. We establish a baseline encoding scheme to translate complex vector data changes into a structured tabular format for classification models. The two crucial components of the encoding scheme include a spatial context encoder to enrich the change representations with surrounding geographic features, and an ENC attribute encoder to represent nuanced attribute-value descriptions of the modified objects. We evaluate the proposed approach across two distinct operational datasets, comprising 1,308 chart pairs containing over 100,000 individual chart modifications. Tuned gradient-boosted trees leveraging the proposed encoding schemes achieve accuracies of 90% and 94% on the two datasets, yielding a 5-7% improvement over default hyperparameterized models trained on encodings without spatial context and attribute embeddings. These results demonstrate the viability of integrating machine learning into operational geospatial pipelines to improve ENC maintenance and enhance maritime safety. Finally, our experiments demonstrate the effectiveness of simple location and spatial aggregation methods, providing a foundation for evaluating more sophisticated spatial representation learning techniques for this application.
2026-09-12
Laura M. Guzmán-Rincón、Ya
人工智能(AI)作为现代数字时代最具影响力的技术之一,依托大数据、算法优化与算力提升的协同进步,已深度融入教育、工业、医疗及日常生活等关键领域。本文系统阐述了人工智能的基本内涵,强调其核心在于赋予机器模拟人类认知、判断与自主学习的能力;重点分析了生成式人工智能与大语言模型的突破性进展,指出其在文本生成、数据分析、语言翻译与逻辑推理等方面展现出显著的自主性与适应性。研究进一步归纳了AI的四大应用方向:一是提升办公与工作效率,通过智能文档生成、数据自动处理与流程优化减少重复性劳动;二是赋能教育与科研,实现个性化学习支持与科研数据分析辅助;三是驱动工业智能化,涵盖设备状态监测、风险预测与安全生产保障;四是重塑日常生活,体现在智能语音交互、个性化内容推荐与自动驾驶等场景。最后,论文探讨了AI在可解释性、伦理治理、数据安全与技术普惠等方面面临的现实挑战,并对其融合多模态感知、强化人机协同、深化垂直领域落地等未来发展趋势进行了展望。
2026-09-12
Mingkuan Feng、Zhengqi Wen、Jianhua Tao
Multimodal large language models (MLLMs) have demonstrated remarkable capabilities by integrating visual and textual understanding within a unified transformer architecture. However, fine-tuning all parameters of these models for visual instruction tuning is computationally expensive and often unnecessary, as the representation requirements for visual and textual tokens diverge significantly in the deeper layers of the network. In this paper, we propose Decoupled Visual Processing (DVP), an efficient training framework that replaces the upper decoder layers of a pretrained LLM with a lightweight, independently trainable single transformer block dedicated exclusively to visual token processing. Specifically, after shared processing through the first half of the decoder layers, visual and textual tokens are split: visual tokens are routed through a newly initialized single transformer block while textual tokens continue through the original frozen decoder layers. The two streams are then concatenated before the language modeling head. During training, only the single transformer block is updated, dramatically reducing the number of trainable parameters. Experiments on the LLaVA-1.5 framework demonstrate that DVP achieves competitive performance on MME, POPE, and ChartQA benchmarks while training only a fraction of the total parameters, suggesting that visual representations in MLLMs can be effectively learned through a decoupled, parameter-efficient pathway.
2026-07-16
Huanghao Yin、Shenkun Xu、Kanle Shi、Junhai Yong、Bin Wang
Text-conditioned image editing has greatly benefited from the advancements in Image Diffusion Models. However, extending these techniques to facial video editing introduces challenges in preserving facial identity throughout the source video and ensuring consistency of the edited subject across frames. In this paper, we introduce DiffMagicFace, a unique video editing framework that integrates two fine-tuned models for text and image control. These models operate concurrently during inference to produce video frames that maintain identity features while seamlessly aligning with the editing semantics. To ensure the consistency of the edited videos, we develop a dataset comprising images showcasing various facial perspectives for each edited subject. The creation of a data set is achieved through rendering techniques and the subsequent application of optimization algorithms. Remarkably, our approach does not depend on video datasets but still delivers high-quality results in both consistency and content. The excellent effect holds even for complex tasks like talking head videos and distinguishing closely related categories. The videos edited using our framework exhibit parity with videos that are made using traditional rendering software. Through comparative analysis with current state-of-the-art methods, our framework demonstrates superior performance in both visual appeal and quantitative metrics.
2026-07-16
Zhe Wu、Hongjin Lu、Junliang Xing、Changhao Zhang、Yuxuan Li、Yin Zhu、Yuhao Yang、Yuheng Jing、Kai Li、Kun Shao、Jianye Hao、Jun Wang、Yuanchun Shi
Building agents that autonomously operate mobile devices has attracted increasing attention. While Vision-Language Models (VLMs) show promise, most existing approaches rely on direct state-to-action mappings, which lack structured reasoning and planning, and thus generalize poorly to novel tasks or unseen UI layouts. We introduce Hi-Agent, a trainable hierarchical vision-language agent for mobile control, featuring a high-level reasoning model and a low-level action model that are jointly optimized. For efficient training, we reformulate multi-step decision-making as a sequence of single-step subgoals and propose a foresight advantage function, which leverages execution feedback from the low-level model to guide high-level optimization. This design alleviates the path explosion issue encountered by Group Relative Policy Optimization (GRPO) in long-horizon tasks and enables stable, critic-free joint training. Hi-Agent achieves a new State-Of-The-Art (SOTA) 87.9% task success rate on the Android-in-the-Wild (AitW) benchmark, significantly outperforming prior methods across three paradigms: prompt-based (AppAgent: 17.7%), supervised (Filtered BC: 54.5%), and reinforcement learning-based (DigiRL: 71.9%). It also demonstrates competitive zero-shot generalization on the ScreenSpot-v2 benchmark. On the more challenging Android-World benchmark, Hi-Agent also scales effectively with larger backbones, showing strong adaptability in high-complexity mobile control scenarios.
2026-07-16
Laura M. Guzmán-Rincón、Ya
人工智能(AI)作为现代数字时代最具影响力的技术之一,依托大数据、算法优化与算力提升的协同进步,已深度融入教育、工业、医疗及日常生活等关键领域。本文系统阐述了人工智能的基本内涵,强调其核心在于赋予机器模拟人类认知、判断与自主学习的能力;重点分析了生成式人工智能与大语言模型的突破性进展,指出其在文本生成、数据分析、语言翻译与逻辑推理等方面展现出显著的自主性与适应性。研究进一步归纳了AI的四大应用方向:一是提升办公与工作效率,通过智能文档生成、数据自动处理与流程优化减少重复性劳动;二是赋能教育与科研,实现个性化学习支持与科研数据分析辅助;三是驱动工业智能化,涵盖设备状态监测、风险预测与安全生产保障;四是重塑日常生活,体现在智能语音交互、个性化内容推荐与自动驾驶等场景。最后,论文探讨了AI在可解释性、伦理治理、数据安全与技术普惠等方面面临的现实挑战,并对其融合多模态感知、强化人机协同、深化垂直领域落地等未来发展趋势进行了展望。
2026-09-12
Mingkuan Feng、Zhengqi Wen、Jianhua Tao
Multimodal large language models (MLLMs) have demonstrated remarkable capabilities by integrating visual and textual understanding within a unified transformer architecture. However, fine-tuning all parameters of these models for visual instruction tuning is computationally expensive and often unnecessary, as the representation requirements for visual and textual tokens diverge significantly in the deeper layers of the network. In this paper, we propose Decoupled Visual Processing (DVP), an efficient training framework that replaces the upper decoder layers of a pretrained LLM with a lightweight, independently trainable single transformer block dedicated exclusively to visual token processing. Specifically, after shared processing through the first half of the decoder layers, visual and textual tokens are split: visual tokens are routed through a newly initialized single transformer block while textual tokens continue through the original frozen decoder layers. The two streams are then concatenated before the language modeling head. During training, only the single transformer block is updated, dramatically reducing the number of trainable parameters. Experiments on the LLaVA-1.5 framework demonstrate that DVP achieves competitive performance on MME, POPE, and ChartQA benchmarks while training only a fraction of the total parameters, suggesting that visual representations in MLLMs can be effectively learned through a decoupled, parameter-efficient pathway.
2026-07-16
Huanghao Yin、Shenkun Xu、Kanle Shi、Junhai Yong、Bin Wang
Text-conditioned image editing has greatly benefited from the advancements in Image Diffusion Models. However, extending these techniques to facial video editing introduces challenges in preserving facial identity throughout the source video and ensuring consistency of the edited subject across frames. In this paper, we introduce DiffMagicFace, a unique video editing framework that integrates two fine-tuned models for text and image control. These models operate concurrently during inference to produce video frames that maintain identity features while seamlessly aligning with the editing semantics. To ensure the consistency of the edited videos, we develop a dataset comprising images showcasing various facial perspectives for each edited subject. The creation of a data set is achieved through rendering techniques and the subsequent application of optimization algorithms. Remarkably, our approach does not depend on video datasets but still delivers high-quality results in both consistency and content. The excellent effect holds even for complex tasks like talking head videos and distinguishing closely related categories. The videos edited using our framework exhibit parity with videos that are made using traditional rendering software. Through comparative analysis with current state-of-the-art methods, our framework demonstrates superior performance in both visual appeal and quantitative metrics.
2026-07-16
Zhe Wu、Hongjin Lu、Junliang Xing、Changhao Zhang、Yuxuan Li、Yin Zhu、Yuhao Yang、Yuheng Jing、Kai Li、Kun Shao、Jianye Hao、Jun Wang、Yuanchun Shi
Building agents that autonomously operate mobile devices has attracted increasing attention. While Vision-Language Models (VLMs) show promise, most existing approaches rely on direct state-to-action mappings, which lack structured reasoning and planning, and thus generalize poorly to novel tasks or unseen UI layouts. We introduce Hi-Agent, a trainable hierarchical vision-language agent for mobile control, featuring a high-level reasoning model and a low-level action model that are jointly optimized. For efficient training, we reformulate multi-step decision-making as a sequence of single-step subgoals and propose a foresight advantage function, which leverages execution feedback from the low-level model to guide high-level optimization. This design alleviates the path explosion issue encountered by Group Relative Policy Optimization (GRPO) in long-horizon tasks and enables stable, critic-free joint training. Hi-Agent achieves a new State-Of-The-Art (SOTA) 87.9% task success rate on the Android-in-the-Wild (AitW) benchmark, significantly outperforming prior methods across three paradigms: prompt-based (AppAgent: 17.7%), supervised (Filtered BC: 54.5%), and reinforcement learning-based (DigiRL: 71.9%). It also demonstrates competitive zero-shot generalization on the ScreenSpot-v2 benchmark. On the more challenging Android-World benchmark, Hi-Agent also scales effectively with larger backbones, showing strong adaptability in high-complexity mobile control scenarios.
2026-07-16