Frontiers of Computer Science

ISSN 2095-2228

ISSN 2095-2236(Online)

CN 10-1014/TP

邮发代号 80-970

2019 Impact Factor: 1.275

   优先出版

合作单位

全文下载排行
一年内发表文章 | 两年内 | 三年内 | 全部 | 最近1个月下载排行 | 最近1年下载排行

当前位置: 最近1年下载排行
Please wait a minute...
选择: 合并摘要 显示/隐藏图片
Sememe knowledge computation: a review of recent advances in application and expansion of sememe knowledge bases
Fanchao QI, Ruobing XIE, Yuan ZANG, Zhiyuan LIU, Maosong SUN
Frontiers of Computer Science    2021, 15 (5): 155327-null.   https://doi.org/10.1007/s11704-020-0002-4
摘要   PDF (526KB)

A sememe is defined as the minimum semantic unit of languages in linguistics. Sememe knowledge bases are built by manually annotating sememes for words and phrases. HowNet is the most well-known sememe knowledge base. It has been extensively utilized in many natural language processing tasks in the era of statistical natural language processing and proven to be effective and helpful to understanding and using languages. In the era of deep learning, although data are thought to be of vital importance, there are some studies working on incorporating sememe knowledge bases like HowNet into neural network models to enhance system performance. Some successful attempts have been made in the tasks including word representation learning, language modeling, semantic composition, etc. In addition, considering the high cost of manual annotation and update for sememe knowledge bases, some work has tried to use machine learning methods to automatically predict sememes for words and phrases to expand sememe knowledge bases. Besides, some studies try to extend HowNet to other languages by automatically predicting sememes for words and phrases in a new language. In this paper, we summarize recent studies on application and expansion of sememe knowledge bases and point out some future directions of research on sememes.

参考文献 | 补充材料 | 相关文章 | 多维度评价
A survey on ensemble learning
Xibin DONG, Zhiwen YU, Wenming CAO, Yifan SHI, Qianli MA
Frontiers of Computer Science    2020, 14 (2): 241-258.   https://doi.org/10.1007/s11704-019-8208-z
摘要   PDF (648KB)

Despite significant successes achieved in knowledge discovery, traditional machine learning methods may fail to obtain satisfactory performances when dealing with complex data, such as imbalanced, high-dimensional, noisy data, etc. The reason behind is that it is difficult for these methods to capture multiple characteristics and underlying structure of data. In this context, it becomes an important topic in the data mining field that how to effectively construct an efficient knowledge discovery and mining model. Ensemble learning, as one research hot spot, aims to integrate data fusion, data modeling, and data mining into a unified framework. Specifically, ensemble learning firstly extracts a set of features with a variety of transformations. Based on these learned features, multiple learning algorithms are utilized to produce weak predictive results. Finally, ensemble learning fuses the informative knowledge from the above results obtained to achieve knowledge discovery and better predictive performance via voting schemes in an adaptive way. In this paper, we review the research progress of the mainstream approaches of ensemble learning and classify them based on different characteristics. In addition, we present challenges and possible research directions for each mainstream approach of ensemble learning, and we also give an extra introduction for the combination of ensemble learning with other machine learning hot spots such as deep learning, reinforcement learning, etc.

参考文献 | 补充材料 | 相关文章 | 多维度评价
A survey on large language model based autonomous agents
Lei WANG, Chen MA, Xueyang FENG, Zeyu ZHANG, Hao YANG, Jingsen ZHANG, Zhiyuan CHEN, Jiakai TANG, Xu CHEN, Yankai LIN, Wayne Xin ZHAO, Zhewei WEI, Jirong WEN
Frontiers of Computer Science    2024, 18 (6): 186345-.   https://doi.org/10.1007/s11704-024-40231-1
摘要   HTML   PDF (4242KB)

Autonomous agents have long been a research focus in academic and industry communities. Previous research often focuses on training agents with limited knowledge within isolated environments, which diverges significantly from human learning processes, and makes the agents hard to achieve human-like decisions. Recently, through the acquisition of vast amounts of Web knowledge, large language models (LLMs) have shown potential in human-level intelligence, leading to a surge in research on LLM-based autonomous agents. In this paper, we present a comprehensive survey of these studies, delivering a systematic review of LLM-based autonomous agents from a holistic perspective. We first discuss the construction of LLM-based autonomous agents, proposing a unified framework that encompasses much of previous work. Then, we present a overview of the diverse applications of LLM-based autonomous agents in social science, natural science, and engineering. Finally, we delve into the evaluation strategies commonly used for LLM-based autonomous agents. Based on the previous studies, we also present several challenges and future directions in this field.

图表 | 参考文献 | 相关文章 | 多维度评价
Identifying useful learnwares via learnable specification
Zhi-Yu SHEN, Ming LI
Frontiers of Computer Science    2025, 19 (9): 199344-null.   https://doi.org/10.1007/s11704-024-40135-0
摘要   HTML   PDF (1468KB)

The learnware paradigm has been proposed as a new manner for reusing models from a market of various well-trained models, which can relieve users’ burden of training a new model from scratch. A learnware consists of a well-trained model and a specification which explains the purpose or specialty of the model without revealing data. By specification matching, the market can identify the most useful learnwares for users’ tasks. Prior art attempted to generate the specification by a reduced kernel mean embedding approach. However, such kind of specification is defined by some pre-designed kernel function, which lacks flexibility. In this paper, we advance a methodology for direct specification learning from data, introducing a novel neural network named SpecNet for this purpose. Our approach accepts unordered datasets as input and subsequently produces specification vectors in a latent space. Notably, the flexibility and efficiency of our learned specifications are underscored by their derivation from diverse tasks, rendering them particularly adept for learnware identification. Empirical studies provide validation for the efficacy of our proposed approach.

图表 | 参考文献 | 补充材料 | 相关文章 | 多维度评价
Top-k probabilistic prevalent co-location mining in spatially uncertain data sets
Lizhen WANG,Jun HAN,Hongmei CHEN,Junli LU
Frontiers of Computer Science    2016, 10 (3): 488-503.   https://doi.org/10.1007/s11704-015-4196-9
摘要   PDF (774KB)

A co-location pattern is a set of spatial features whose instances frequently appear in a spatial neighborhood. This paper efficiently mines the top-k probabilistic prevalent co-locations over spatially uncertain data sets and makes the following contributions: 1) the concept of the top-k probabilistic prevalent co-locations based on a possible world model is defined; 2) a framework for discovering the top-k probabilistic prevalent co-locations is set up; 3) a matrix method is proposed to improve the computation of the prevalence probability of a top-k candidate, and two pruning rules of the matrix block are given to accelerate the search for exact solutions; 4) a polynomial matrix is developed to further speed up the top-k candidate refinement process; 5) an approximate algorithm with compensation factor is introduced so that relatively large quantity of data can be processed quickly. The efficiency of our proposed algorithms as well as the accuracy of the approximation algorithms is evaluated with an extensive set of experiments using both synthetic and real uncertain data sets.

参考文献 | 补充材料 | 相关文章 | 多维度评价
Distributed top-k similarity query on big trajectory streams
Zhigang ZHANG, Xiaodong QI, Yilin WANG, Cheqing JIN, Jiali MAO, Aoying ZHOU
Frontiers of Computer Science    2019, 13 (3): 647-664.   https://doi.org/10.1007/s11704-018-7234-6
摘要   PDF (742KB)

Recently, big trajectory data streams are generated in distributed environmentswith the popularity of smartphones and other mobile devices. Distributed top-k similarity query, which finds k trajectories that are most similar to a given query trajectory from all remote sites, is critical in this field. The key challenge in such a query is how to reduce the communication cost due to the limited network bandwidth resource. Although this query can be solved by sending the query trajectory to all the remote sites, in which the pairwise similarities are computed precisely. However, the overall cost, O(n · m), is huge when n or m is huge, where n is the size of query trajectory and m is the number of remote sites. Fortunately, there are some cheap ways to estimate pairwise similarity, which filter some trajectories in advance without precise computation. In order to overcome the challenge in this query, we devise two general frameworks, into which concrete distance measures can be plugged. The former one uses two bounds (the upper and lower bound), while the latter one only uses the lower bound. Moreover, we introduce detailed implementations of two representative distance measures, Euclidean and DTW distance, after inferring the lower and upper bound for the former framework and the lower bound for the latter one. Theoretical analysis and extensive experiments on real-world datasets evaluate the efficiency of proposed methods.

参考文献 | 补充材料 | 相关文章 | 多维度评价
Co-occurrence prediction in a large location-based social network
Rong-Hua LI, Jianquan LIU, Jeffrey Xu YU, Hanxiong CHEN, Hiroyuki KITAGAWA
Frontiers of Computer Science    2013, 7 (2): 185-194.   https://doi.org/10.1007/s11704-013-3902-8
摘要   HTML   PDF (543KB)

Location-based social network (LBSN) is at the forefront of emerging trends in social network services (SNS) since the users in LBSN are allowed to “check-in” the places (locations) when they visit them. The accurate geographical and temporal information of these check-in actions are provided by the end-user GPS-enabled mobile devices, and recorded by the LBSN system. In this paper, we analyze and mine a big LBSN data, Gowalla, collected by us. First, we investigate the relationship between the spatio-temporal cooccurrences and social ties, and the results show that the cooccurrences are strongly correlative with the social ties. Second, we present a study of predicting two users whether or not they will meet (co-occur) at a place in a given future time, by exploring their check-in habits. In particular, we first introduce two new concepts, bag-of-location and bag-of-time-lag, to characterize user’s check-in habits. Based on such bag representations, we define a similarity metric called habits similarity to measure the similarity between two users’ check-in habits. Then we propose a machine learning formula for predicting co-occurrence based on the social ties and habits similarities. Finally, we conduct extensive experiments on our dataset, and the results demonstrate the effectiveness of the proposed method.

参考文献 | 相关文章 | 多维度评价
Inforence: effective fault localization based on information-theoretic analysis and statistical causal inference
Farid FEYZI, Saeed PARSA
Frontiers of Computer Science    2019, 13 (4): 735-759.   https://doi.org/10.1007/s11704-017-6512-z
摘要   PDF (1378KB)

In this paper, a novel approach, Inforence, is proposed to isolate the suspicious codes that likely contain faults. Inforence employs a feature selection method, based on mutual information, to identify those bug-related statements that may cause the program to fail. Because the majority of a program faults may be revealed as undesired joint effect of the program statements on each other and on program termination state, unlike the state-of-the-art methods, Inforence tries to identify and select groups of interdependent statements which altogether may affect the program failure. The interdependence amongst the statements is measured according to their mutual effect on each other and on the program termination state. To provide the context of failure, the selected bug-related statements are chained to each other, considering the program static structure. Eventually, the resultant causeeffect chains are ranked according to their combined causal effect on program failure. To validate Inforence, the results of our experimentswith seven sets of programs include Siemens suite, gzip, grep, sed, space, make and bash are presented. The experimental results are then compared with those provided by different fault localization techniques for the both single-fault and multi-fault programs. The experimental results prove the outperformance of the proposed method compared to the state-of-the-art techniques.

参考文献 | 补充材料 | 相关文章 | 多维度评价
Gria: an efficient deterministic concurrency control protocol
Xinyuan WANG, Yun PENG, Hejiao HUANG
Frontiers of Computer Science    2024, 18 (4): 184204-null.   https://doi.org/10.1007/s11704-023-2605-z
摘要   HTML   PDF (13477KB)

Deterministic databases are able to reduce coordination costs in a replication. This property has fostered a significant interest in the design of efficient deterministic concurrency control protocols. However, the state-of-the-art deterministic concurrency control protocol Aria has three issues. First, it is impractical to configure a suitable batch size when the read-write set is unknown. Second, Aria running in low-concurrency scenarios, e.g., a single-thread scenario, suffers from the same conflicts as running in high-concurrency scenarios. Third, the single-version schema brings write-after-write conflicts.

To address these issues, we propose Gria, an efficient deterministic concurrency control protocol. Gria has the following properties. First, the batch size of Gria is auto-scaling. Second, Gria’s conflict probability in low-concurrency scenarios is lower than that in high-concurrency scenarios. Third, Gria has no write-after-write conflicts by adopting a multi-version structure. To further reduce conflicts, we propose two optimizations: a reordering mechanism as well as a rechecking strategy. The evaluation result on two popular benchmarks shows that Gria outperforms Aria by 13x.

图表 | 参考文献 | 补充材料 | 相关文章 | 多维度评价
ncRNA2MetS v2.0: a manually curated database for metabolic syndrome-associated ncRNAs
Dengju YAO, Zhanhe LI, Xiaojuan ZHAN, Zibin ZHOU, Hao LIANG
Frontiers of Computer Science    2025, 19 (6): 196911-null.   https://doi.org/10.1007/s11704-024-40709-y
摘要   HTML   PDF (714KB)
图表 | 参考文献 | 补充材料 | 相关文章 | 多维度评价
Locality-constrained framework for face alignment
Jie ZHANG, Xiaowei ZHAO, Meina KAN, Shiguang SHAN, Xiujuan CHAI, Xilin CHEN
Frontiers of Computer Science    2019, 13 (4): 789-801.   https://doi.org/10.1007/s11704-018-6617-z
摘要   PDF (711KB)

Although the conventional active appearance model (AAM) has achieved some success for face alignment, it still suffers from the generalization problem when be applied to unseen subjects and images. To deal with the generalization problem of AAM, we first reformulate the original AAM as sparsity-regularized AAM, which can achieve more compact/better shape and appearance priors by selecting nearest neighbors as the bases of the shape and appearance model. To speed up the fitting procedure, the sparsity in sparsity-regularized AAM is approximated by using the locality (i.e., K-nearest neighbor), and thus inducing the locality-constrained active appearancemodel (LC-AAM). The LC-AAM solves a constrained AAM-like fitting problem with the K-nearest neighbors as the bases of shape and appearance model. To alleviate the adverse influence of inaccurate K-nearest neighbor results, the locality constraint is further embedded in the discriminative fitting method denoted as LC-DFM, which can find better K-nearest neighbor results by employing shape-indexed feature, and can also tolerate some inaccurate neighbors benefited from the regression model rather than the generative model in AAM. Extensive experiments on several datasets demonstrate that our methods outperform the state-of-the-arts in both detection accuracy and generalization ability.

参考文献 | 补充材料 | 相关文章 | 多维度评价
Vehicle color recognition based on smooth modulation neural network with multi-scale feature fusion
Mingdi HU, Long BAI, Jiulun FAN, Sirui ZHAO, Enhong CHEN
Frontiers of Computer Science    2023, 17 (3): 173321-null.   https://doi.org/10.1007/s11704-022-1389-x
摘要   HTML   PDF (8165KB)

Vehicle Color Recognition (VCR) plays a vital role in intelligent traffic management and criminal investigation assistance. However, the existing vehicle color datasets only cover 13 classes, which can not meet the current actual demand. Besides, although lots of efforts are devoted to VCR, they suffer from the problem of class imbalance in datasets. To address these challenges, in this paper, we propose a novel VCR method based on Smooth Modulation Neural Network with Multi-Scale Feature Fusion (SMNN-MSFF). Specifically, to construct the benchmark of model training and evaluation, we first present a new VCR dataset with 24 vehicle classes, Vehicle Color-24, consisting of 10091 vehicle images from a 100-hour urban road surveillance video. Then, to tackle the problem of long-tail distribution and improve the recognition performance, we propose the SMNN-MSFF model with multi-scale feature fusion and smooth modulation. The former aims to extract feature information from local to global, and the latter could increase the loss of the images of tail class instances for training with class-imbalance. Finally, comprehensive experimental evaluation on Vehicle Color-24 and previously three representative datasets demonstrate that our proposed SMNN-MSFF outperformed state-of-the-art VCR methods. And extensive ablation studies also demonstrate that each module of our method is effective, especially, the smooth modulation efficiently help feature learning of the minority or tail classes. Vehicle Color-24 and the code of SMNN-MSFF are publicly available and can contact the author to obtain.

图表 | 参考文献 | 补充材料 | 相关文章 | 多维度评价
System architecture for high-performance permissioned blockchains
Libo FENG, Hui ZHANG, Wei-Tek TSAI, Simeng SUN
Frontiers of Computer Science    2019, 13 (6): 1151-1165.   https://doi.org/10.1007/s11704-018-6345-4
摘要   PDF (695KB)

Blockchain(BC), as an emerging distributed database technology with advanced security and reliability, has attracted much attention from experts who devoted to e-finance, intellectual property protection, the internet of things (IoT) and so forth. However, the inefficient transaction processing speed, which hinders the BC’s widespread, has not been well tackled yet. In this paper, we propose a novel architecture, called Dual-Channel Parallel Broadcast model (DCPB), which could address such a problem to a greater extent by using three methods which are dual communication channels, parallel pipeline processing and block broadcast strategy. In the dual-channel model, one channel processes transactions, and the other engages in the execution of BFT. The parallel pipeline processing allows the system to operate asynchronously. The block generation strategy improves the efficiency and speed of processing. Extensive experiments have been applied to BeihangChain, a simplified prototype for BC system, illustrates that its transaction processing speed could be improved to 16K transaction per second which could well supportmany real-world scenarios such as BC-based energy trading system andMicro-film copyright trading system in CCTV.

参考文献 | 补充材料 | 相关文章 | 多维度评价
Towards efficient and effective unlearning of large language models for recommendation
Hangyu WANG, Jianghao LIN, Bo CHEN, Yang YANG, Ruiming TANG, Weinan ZHANG, Yong YU
Frontiers of Computer Science    2025, 19 (3): 193327-null.   https://doi.org/10.1007/s11704-024-40044-2
摘要   HTML   PDF (406KB)
图表 | 参考文献 | 补充材料 | 相关文章 | 多维度评价
BIFER: a biphasic trace filter approach to scalable prediction of concurrency errors
Xi CHANG,Zhuo ZHANG,Peng ZHANG,Jianxin XUE,Jianjun ZHAO
Frontiers of Computer Science    2015, 9 (6): 944-955.   https://doi.org/10.1007/s11704-015-4334-4
摘要   PDF (1180KB)

Predictive trace analysis (PTA), a static trace analysis technique for concurrent programs, can offer powerful capability support for finding concurrency errors unseen in a previous program execution. Existing PTA techniques always face considerable challenges in scaling to large traces which contain numerous critical events. One main reason is that an analyzed trace includes not only redundant memory accessing events and threads that cannot contribute to discovering any additional errors different from the found candidate ones, but also many residual synchronization events which still affect PTA to check whether these candidate ones are feasible or not even after removing the redundant events. Removing them from the trace can significantly improve the scalability of PTA without affecting the quality of the PTA results. In this paper, we propose a biphasic trace filter approach, BIFER in short, to filter these redundant events and residual events for improving the scalability of PTA to expose general concurrency errors. In addition, we design a model which indicates the lock history and the happens-before history of each thread with two kinds of ways to achieve the efficient filtering. We implement a prototypical tool BIFER for Java programs on the basis of a predictive trace analysis framework. Experiments show that BIFER can improve the scalability of PTA during the process of analyzing all of the traces.

参考文献 | 补充材料 | 相关文章 | 多维度评价
Single depth image 3D face reconstruction via domain adaptive learning
Xiaoxu CAI, Jianwen LOU, Jiajun BU, Junyu DONG, Haishuai WANG, Hui YU
Frontiers of Computer Science    2024, 18 (1): 181342-.   https://doi.org/10.1007/s11704-023-3541-7
摘要   HTML   PDF (1600KB)
图表 | 参考文献 | 补充材料 | 相关文章 | 多维度评价
Adaptive network combination for single-image reflection removal: a domain generalization perspective
Ming LIU, Jianan PAN, Zifei YAN, Wangmeng ZUO, Lei ZHANG
Frontiers of Computer Science    2025, 19 (1): 191703-null.   https://doi.org/10.1007/s11704-024-3582-6
摘要   HTML   PDF (1253KB)
图表 | 参考文献 | 补充材料 | 相关文章 | 多维度评价
A PTS-PGATS based approach for data-intensive scheduling in data grids
Kenli LI, Zhao TONG, Dan LIU, Teklay TESFAZGHI, Xiangke LIAO
Frontiers of Computer Science in China    2011, 5 (4): 513-525.   https://doi.org/10.1007/s11704-011-0970-5
摘要   HTML   PDF (521KB)

Grid computing is the combination of computer resources in a loosely coupled, heterogeneous, and geographically dispersed environment. Grid data are the data used in grid computing, which consists of large-scale data-intensive applications, producing and consuming huge amounts of data, distributed across a large number of machines. Data grid computing composes sets of independent tasks each of which require massive distributed data sets that may each be replicated on different resources. To reduce the completion time of the application and improve the performance of the grid, appropriate computing resources should be selected to execute the tasks and appropriate storage resources selected to serve the files required by the tasks. So the problem can be broken into two sub-problems: selection of storage resources and assignment of tasks to computing resources. This paper proposes a scheduler, which is broken into three parts that can run in parallel and uses both parallel tabu search and a parallel genetic algorithm. Finally, the proposed algorithm is evaluated by comparing it with other related algorithms, which target minimizing makespan. Simulation results show that the proposed approach can be a good choice for scheduling large data grid applications.

图表 | 参考文献 | 相关文章 | 多维度评价
Clustered Reinforcement Learning
Xiao MA, Shen-Yi ZHAO, Zhao-Heng YIN, Wu-Jun LI
Frontiers of Computer Science    2025, 19 (4): 194313-.   https://doi.org/10.1007/s11704-024-3194-1
摘要   HTML   PDF (2856KB)

Exploration strategy design is a challenging problem in reinforcement learning (RL), especially when the environment contains a large state space or sparse rewards. During exploration, the agent tries to discover unexplored (novel) areas or high reward (quality) areas. Most existing methods perform exploration by only utilizing the novelty of states. The novelty and quality in the neighboring area of the current state have not been well utilized to simultaneously guide the agent’s exploration. To address this problem, this paper proposes a novel RL framework, called clustered reinforcement learning (CRL), for efficient exploration in RL. CRL adopts clustering to divide the collected states into several clusters, based on which a bonus reward reflecting both novelty and quality in the neighboring area (cluster) of the current state is given to the agent. CRL leverages these bonus rewards to guide the agent to perform efficient exploration. Moreover, CRL can be combined with existing exploration strategies to improve their performance, as the bonus rewards employed by these existing exploration strategies solely capture the novelty of states. Experiments on four continuous control tasks and six hard-exploration Atari-2600 games show that our method can outperform other state-of-the-art methods to achieve the best performance.

图表 | 参考文献 | 补充材料 | 相关文章 | 多维度评价
Foundation model enhanced derivative-free cognitive diagnosis
Mingjia LI, Hong QIAN, Jinglan LV, Mengliang HE, Wei ZHANG, Aimin ZHOU
Frontiers of Computer Science    2025, 19 (1): 191318-null.   https://doi.org/10.1007/s11704-024-40029-1
摘要   HTML   PDF (934KB)
图表 | 参考文献 | 相关文章 | 多维度评价
Deterministic streaming algorithms for non-monotone submodular maximization
Xiaoming SUN, Jialin ZHANG, Shuo ZHANG
Frontiers of Computer Science    2025, 19 (6): 196404-null.   https://doi.org/10.1007/s11704-024-40266-4
摘要   HTML   PDF (1672KB)

Submodular maximization is a significant area of interest in combinatorial optimization. It has various real-world applications. In recent years, streaming algorithms for submodular maximization have gained attention, allowing real-time processing of large data sets by examining each piece of data only once. However, most of the current state-of-the-art algorithms are only applicable to monotone submodular maximization. There are still significant gaps in the approximation ratios between monotone and non-monotone objective functions.

In this paper, we propose a streaming algorithm framework for non-monotone submodular maximization and use this framework to design deterministic streaming algorithms for the d-knapsack constraint and the knapsack constraint. Our 1-pass streaming algorithm for the d-knapsack constraint has a 14(d+1)ϵ approximation ratio, using O(B~logB~ϵ) memory, and O(logB~ϵ) query time per element, where B~=min(n,b) is the maximum number of elements that the knapsack can store. As a special case of the d-knapsack constraint, we have the 1-pass streaming algorithm with a 1/8ϵ approximation ratio to the knapsack constraint. To our knowledge, there is currently no streaming algorithm for this constraint when the objective function is non-monotone, even when d = 1. In addition, we propose a multi-pass streaming algorithm with 1/6ϵ approximation, which stores O(B~) elements.

图表 | 参考文献 | 补充材料 | 相关文章 | 多维度评价
Graph foundation model
Chuan SHI, Junze CHEN, Jiawei LIU, Cheng YANG
Frontiers of Computer Science    2024, 18 (6): 186355-null.   https://doi.org/10.1007/s11704-024-40046-0
摘要   HTML   PDF (664KB)
图表 | 参考文献 | 补充材料 | 相关文章 | 多维度评价
ROS package search for robot software development: a knowledge graph-based approach
Shuo WANG, Xinjun MAO, Shuo YANG, Menghan WU, Zhang ZHANG
Frontiers of Computer Science    2025, 19 (6): 196320-null.   https://doi.org/10.1007/s11704-024-3660-9
摘要   HTML   PDF (5039KB)

In recent years, ROS (Robot Operating System) packages have become increasingly popular as a type of software artifact that can be effectively reused in robotic software development. Indeed, finding suitable ROS packages that closely match the software’s functional requirements from the vast number of available packages is a nontrivial task using current search methods. The traditional search methods for ROS packages often involve inputting keywords related to robotic tasks into general-purpose search engines (e.g., Google) or code hosting platforms (e.g., GitHub) to obtain approximate results of all potentially suitable ROS packages. However, the accuracy of these search methods remains relatively low because the task-related keywords may not precisely match the functionalities offered by the ROS packages. To improve the search accuracy of ROS packages, this paper presents a novel semantic-based search approach that relies on the semantic-level ROS Package Knowledge Graph (RPKG) to automatically retrieve the most suitable ROS packages. Firstly, to construct the RPKG, we employ multi-dimensional feature extraction techniques to extract semantic concepts, including code file name, category, hardware device, characteristics, and function, from the dataset of ROS package text descriptions. The semantic features extracted from this process result in a substantial number of entities (32,294) and relationships (54,698). Subsequently, we create a robot domain-specific small corpus and further fine-tune a pre-trained language model, BERT-ROS, to generate embeddings that effectively represent the semantics of the extracted features. These embeddings play a crucial role in facilitating semantic-level understanding and comparisons during the ROS package search process within the RPKG. Secondly, we introduce a novel semantic matching-based search algorithm that incorporates the weighted similarities of multiple features from user search queries, which searches out more accurate ROS packages than the traditional keyword search method. To validate the enhanced accuracy of ROS package searching, we conduct comparative case studies between our semantic-based search approach and four baseline search approaches: ROS Index, GitHub, Google, and ChatGPT. The experiment results demonstrate that our approach achieves higher accuracy in terms of ROS package searching, outperforming the other approaches by at least 21% from 5 levels, including top1, top5, top10, top15, and top20.

图表 | 参考文献 | 补充材料 | 相关文章 | 多维度评价
Toward secure and private service discovery anywhere anytime
Feng ZHU, Anish BIVALKAR, Abdullah DEMIR, Yue LU, Chockalingam CHIDAMBARM, Matt MUTKA,
Front. Comput. Sci.    2010, 4 (3): 311-323.   https://doi.org/10.1007/s11704-010-0389-4
摘要   PDF (297KB)
With the advances in and convergence of Internet technologies, embedded computers, and wireless communication, computing devices have become part of our daily life. Hand-held devices and sensors with wireless connections create opportunities for many new nomadic applications. Service discovery is an essential component for cognitive science to discover existing network services just-in-time. Unlike many other approaches, we propose a service discovery model supporting nomadic users and services in public environments. Our model emphasizes secure and private service discovery in such environments. Location sensing is integrated for location dependent service discovery and is used to lessen service discovery network infrastructure requirements. We analyze the system performance and show our formal verification of the protocols. Our implementation shows that our model is feasible.
参考文献 | 相关文章 | 多维度评价
RDF partitioning for scalable SPARQL query processing
Xiaoyan WANG,Tao YANG,Jinchuan CHEN,Long HE,Xiaoyong DU
Frontiers of Computer Science    2015, 9 (6): 919-933.   https://doi.org/10.1007/s11704-015-4104-3
摘要   PDF (1014KB)

The volume of RDF data increases dramatically within recent years, while cloud computing platforms like Hadoop are supposed to be a good choice for processing queries over huge data sets for their wonderful scalability. Previous work on evaluating SPARQL queries with Hadoop mainly focus on reducing the number of joins through careful split of HDFS files and algorithms for generating Map/Reduce jobs. However, the way of partitioning RDF data could also affect system performance. Specifically, a good partitioning solution would greatly reduce or even totally avoid cross-node joins, and significantly cut down the cost in query evaluation. Based on HadoopDB, this work processes SPARQL queries in a hybrid architecture, where Map/Reduce takes charge of the computing tasks, and RDF query engines like RDF-3X store the data and execute join operations. According to the analysis of query workloads, this work proposes a novel algorithm for automatically partitioning RDF data and an approximate solution to physically place the partitions in order to reduce data redundancy. It also discusses how to make a good trade-off between query evaluation efficiency and data redundancy. All of these proposed approaches have been evaluated by extensive experiments over large RDF data sets.

参考文献 | 补充材料 | 相关文章 | 多维度评价
Visual tracking using discriminative representation with 2 regularization
Haijun WANG, Hongjuan GE
Frontiers of Computer Science    2019, 13 (1): 199-211.   https://doi.org/10.1007/s11704-017-6434-9
摘要   PDF (1639KB)

In this paper, we propose a novel visual tracking method using a discriminative representation under a Bayesian framework. First, we exploit the histogram of gradient (HOG) to generate the texture features of the target templates and candidates. Second, we introduce a novel discriminative representation and 2-regularized least squares method to solve the proposed representation model. The proposed model has a closed-form solution and very high computational efficiency. Third, a novel likelihood function and an update scheme considering the occlusion factor are adopted to improve the tracking performance of our proposed method. Both qualitative and quantitative evaluations on 15 challenging video sequences demonstrate that our method can achieve more robust tracking results in terms of the overlap rate and center location error.

参考文献 | 补充材料 | 相关文章 | 多维度评价
Joint salient object detection and existence prediction
Huaizu JIANG, Ming-Ming CHENG, Shi-Jie LI, Ali BORJI, Jingdong WANG
Frontiers of Computer Science    2019, 13 (4): 778-788.   https://doi.org/10.1007/s11704-017-6613-8
摘要   PDF (565KB)

Recent advances in supervised salient object detection modeling has resulted in significant performance improvements on benchmark datasets. However, most of the existing salient object detection models assume that at least one salient object exists in the input image. Such an assumption often leads to less appealing saliencymaps on the background images with no salient object at all. Therefore, handling those cases can reduce the false positive rate of a model. In this paper, we propose a supervised learning approach for jointly addressing the salient object detection and existence prediction problems. Given a set of background-only images and images with salient objects, as well as their salient object annotations, we adopt the structural SVM framework and formulate the two problems jointly in a single integrated objective function: saliency labels of superpixels are involved in a classification term conditioned on the salient object existence variable, which in turn depends on both global image and regional saliency features and saliency labels assignments. The loss function also considers both image-level and regionlevel mis-classifications. Extensive evaluation on benchmark datasets validate the effectiveness of our proposed joint approach compared to the baseline and state-of-the-art models.

参考文献 | 补充材料 | 相关文章 | 多维度评价
ARCHER: a ReRAM-based accelerator for compressed recommendation systems
Xinyang SHEN, Xiaofei LIAO, Long ZHENG, Yu HUANG, Dan CHEN, Hai JIN
Frontiers of Computer Science    2024, 18 (5): 185607-.   https://doi.org/10.1007/s11704-023-3397-x
摘要   HTML   PDF (7144KB)

Modern recommendation systems are widely used in modern data centers. The random and sparse embedding lookup operations are the main performance bottleneck for processing recommendation systems on traditional platforms as they induce abundant data movements between computing units and memory. ReRAM-based processing-in-memory (PIM) can resolve this problem by processing embedding vectors where they are stored. However, the embedding table can easily exceed the capacity limit of a monolithic ReRAM-based PIM chip, which induces off-chip accesses that may offset the PIM profits. Therefore, we deploy the decomposed model on-chip and leverage the high computing efficiency of ReRAM to compensate for the decompression performance loss. In this paper, we propose ARCHER, a ReRAM-based PIM architecture that implements fully on-chip recommendations under resource constraints. First, we make a full analysis of the computation pattern and access pattern on the decomposed table. Based on the computation pattern, we unify the operations of each layer of the decomposed model in multiply-and-accumulate operations. Based on the access observation, we propose a hierarchical mapping schema and a specialized hardware design to maximize resource utilization. Under the unified computation and mapping strategy, we can coordinate the inter-processing elements pipeline. The evaluation shows that ARCHER outperforms the state-of-the-art GPU-based DLRM system, the state-of-the-art near-memory processing recommendation system RecNMP, and the ReRAM-based recommendation accelerator REREC by 15.79×, 2.21×, and 1.21 × in terms of performance and 56.06 ×, 6.45×, and 1.71 × in terms of energy savings, respectively.

图表 | 参考文献 | 补充材料 | 相关文章 | 多维度评价
Revisiting multi-dimensional classification from a dimension-wise perspective
Yi SHI, Hanjia YE, Dongliang MAN, Xiaoxu HAN, Dechuan ZHAN, Yuan JIANG
Frontiers of Computer Science    2025, 19 (1): 191304-.   https://doi.org/10.1007/s11704-023-3272-9
摘要   HTML   PDF (13655KB)

Real-world objects exhibit intricate semantic properties that can be characterized from a multitude of perspectives, which necessitates the development of a model capable of discerning multiple patterns within data, while concurrently predicting several Labeling Dimensions (LDs) — a task known as Multi-dimensional Classification (MDC). While the class imbalance issue has been extensively investigated within the multi-class paradigm, its study in the MDC context has been limited due to the imbalance shift phenomenon. A sample’s classification as a minor or major class instance becomes ambiguous when it belongs to a minor class in one LD and a major class in another. Previous MDC methodologies predominantly emphasized instance-wise criteria, neglecting prediction capabilities from a dimension aspect, i.e., the average classification performance across LDs. We assert the significance of dimension-wise metrics in real-world MDC applications and introduce two such metrics. Furthermore, we observe imbalanced class distributions within each LD and propose a novel Imbalance-Aware fusion Model (IMAM) for addressing the MDC problem. Specifically, we first decompose the task into multiple multi-class classification problems, creating imbalance-aware deep models for each LD separately. This straightforward method performs well across LDs without sacrificing performance in instance-wise criteria. Subsequently, we employ LD-wise models as multiple teachers and transfer their knowledge across all LDs to a unified student model. Experimental results on several real-world datasets demonstrate that our IMAM approach excels in both instance-wise evaluations and the proposed dimension-wise metrics.

图表 | 参考文献 | 补充材料 | 相关文章 | 多维度评价
WPIA: accelerating DNN warm-up in Web browsers by precompiling WebGL programs
Deyu TIAN, Yun MA, Yudong HAN, Qi YANG, Haochen YANG, Gang HUANG
Frontiers of Computer Science    2024, 18 (6): 186211-null.   https://doi.org/10.1007/s11704-024-40066-w
摘要   HTML   PDF (2135KB)
图表 | 参考文献 | 补充材料 | 相关文章 | 多维度评价