Articles in this Volume

Research Article Open Access
Large Language Model self-reflection: advances, limitations and future directions
Large Language Model (LLM) has achieved great success in many fields over recent years, demonstrating a promising future. Nevertheless, hallucination remains a major impediment that limits the usability and reliability of LLM. In order to mitigate hallucination, researchers have proposed various approaches, among which self-reflection stands out. While the initial optimism about pure LLM self-reflection slowly fades away, its methodologies are becoming more complicated and interweaved with other approaches. Due to fact that the boundaries between self-reflection and many of its synonyms are ambiguous, and that researchers keep coining new concepts and terminologies that may overlap and intertwine intricately with each other, this paper aims to provide a clear definition of self-reflection and some frequently used terminologies. In addition, given that currently there is no systematic work sorting out papers in this field, this paper fills this gap by classifying different self-reflection methodologies and presenting paradigmatic researches in each category. By the end, this paper discusses the possible limitations and directions of future development in this field, predicting that self-reflection might evolve into a component in hybrid LLM training approaches that is critical to hallucination alleviation.
Show more
Read Article PDF
Cite
Research Article Open Access
A reliable rolling bearing fault diagnosis method based on Titan
Article thumbnail
The accuracy of fault diagnosis for rolling bearings degrades sharply when operating conditions shift. Existing high-precision classifiers often experience a drop of over 50% in predictive accuracy when speed or load fluctuates, which seriously jeopardizes the reliability of industrial equipment health monitoring. Titan, a recently proposed architecture for long-context language modelling, tackles a similar challenge of maintaining performance across varying contexts. TitanDiag adapts this mechanism to fault diagnosis. The underlying rationale is that a persistent memory accumulates evidence across operating conditions and stabilizes predictions when the current segment alone is ambiguous. The architecture places Titan's dual-path memory (a long-term store gated by surprise plus a short-term FIFO buffer) inside a Transformer encoder. The multi-view front-end provides three complementary representations for every vibration segment, namely the raw waveform, the Fourier magnitude spectrum, and the continuous wavelet transform scalogram. At inference, Monte Carlo dropout produces per-prediction uncertainty scores that align naturally with Titan's surprise metric. On the CWRU and PU bearing benchmarks, TitanDiag attains 99.25% accuracy on the challenging PU-C2 low-speed condition, where TimeMachine and TSCMamba drop to 41.68% and 60.20%, respectively. The mean error-detection AUROC reaches 0.9668, well above the best baseline of 0.9391, demonstrating that the memory-driven variance inflation produces uncertainty estimates that are closely aligned with actual misclassification patterns.
Show more
Read Article PDF
Cite
Research Article Open Access
Key technologies and recent advances in online monitoring of welding quality
Article thumbnail
Online monitoring of welding quality is a key technology in intelligent manufacturing and process closed-loop control. Its core essence is the paradigm shift from passive post-weld inspection to in-process active perception, real-time diagnosis, and closed-loop feedback. Currently, technological development in this field advances along three main paths: first, vision-based detection technology, which provides process information by analyzing weld pool morphology, weld geometry, and surface defects; second, multi-modal sensor-based detection technology, which utilizes acoustic emission, arc/spectrum, thermal imaging, and ultrasonic methods to reflect the physical state and internal changes of the welding process; third, multi-source information fusion technology, which integrates the complementary advantages of multi-source heterogeneous sensing data to enhance diagnostic accuracy and robustness. This paper systematically reviews online monitoring technologies for welding quality: first, it analyzes single-modal detection methods based on vision and multi-sensing; then, taking fusion hierarchy as the dimension, it organizes the technical framework of data-level, feature-level, and decision-level fusion, with emphasis on the frontier progress of deep learning-driven feature-level fusion; finally, it envisions future development trends and key challenges from three dimensions—advanced sensing innovation, large-model enablement, and digital twin.
Show more
Read Article PDF
Cite
Research Article Open Access
Sustainability assessment and improvement of generative AI in industrial product appearance design
As the consumer market continues to expand, relatively homogeneous approaches to industrial product appearance design can no longer meet consumers' changing needs. This study aims to examine the sustainability assessment and improvement of generative AI in industrial product appearance design, with the goal of adapting to evolving consumer demands while promoting continuous innovation. An analysis of the requirements for industrial product appearance design shows that the application of generative AI offers significant advantages. It can not only reduce production costs but also enhance the market appeal of products. Under these conditions, the application of generative AI in industrial product appearance design is becoming increasingly widespread and demonstrates promising prospects for future development.
Show more
Read Article PDF
Cite
Research Article Open Access
Design and operation of simulation experiments based on an AI-empowered predation model
Article thumbnail
Predation is a key biological factor influencing population dynamics, yet traditional teaching methods often struggle to visually illustrate its synchronous periodicity and causal cycles. By leveraging AI to build a predation model, we have developed a predator-prey simulation system. The web-based platform employs a three-pane layout to integrate ecological simulations, data charts, and functional controls. It supports customizable parameters, allows for comparisons between J-shaped and S-shaped growth curves, and includes a CSV data export feature, all of which help foster students' interdisciplinary literacy.
Show more
Read Article PDF
Cite
Research Article Open Access
Research on traffic sign image enhancement model under severe weather conditions
Article thumbnail
With the development of intelligent transportation and intelligent connected vehicles, the accuracy of traffic sign recognition directly affects intelligent driving decision-making and driving safety. Nevertheless, traffic sign recognition suffers from issues such as blurred images, reduced contrast and difficult feature extraction under severe weather including rainy, foggy and nighttime conditions. To tackle the above problems, this paper proposes a joint model integrating image enhancement and recognition. A coding-attention-decoding structured image enhancement module is adopted to strengthen the edges of traffic signs against image quality degradation caused by severe weather. Meanwhile, a loss optimization model based on YOLO is established to reduce the missed detection rate and false detection rate of target recognition. Finally, severe weather data are simulated based on public datasets, and experiments verify the effectiveness of the proposed joint model, which provides effective support for the subsequent technological development of intelligent vehicles.
Show more
Read Article PDF
Cite
Research Article Open Access
Effect of base perforations on natural convection heat dissipation of a finned heat sink
Article thumbnail
Conventional finned heat sinks for highly integrated electronics suffer from blocked airflow, extensive flow dead zones at fin roots and high thermal resistance under natural convection cooling. To simultaneously improve heat dissipation and reduce heat sink weight, this study proposes machining through holes on the heat sink base. Steady laminar finite volume numerical simulations are carried out to compare three perforation shapes: cylindrical, positive frustoconical and inverted frustoconical holes. We analyze how the distance between perforations and heat source influences thermal resistance and flow field, and clarify the heat transfer enhancement mechanism induced by the chimney effect of through holes. Numerical data show that cylindrical-hole heat sinks achieve a 17.44% lower thermal resistance and 15.79% lighter mass, alongside a 43.59% higher mass-specific heat transfer coefficient, compared with non-porous heat sinks. Cylindrical holes maintain uniform and stable airflow and outperform two frustoconical structures. Thermal resistance reaches the minimum value of 13.76 K/W when holes are placed close to the heat source. Base perforations form vertical air passages to suppress recirculation zones at fin roots, providing theoretical guidance for designing natural convection heat sink of high-heat-flux electronic components.
Show more
Read Article PDF
Cite
Research Article Open Access
How many orientations do you need? A dual-polarisation few-shot benchmark for thin-section mineral classification on MUMDMC2025
Article thumbnail
Automated mineral identification in thin section promises to relieve a slow and subjective bottleneck in petrography, yet most deep-learning classifiers use a single polarisation mode recorded at a single stage orientation, discarding information that human petrographers actively exploit. The recently released MUMDMC2025 dataset pairs plane-polarised (PPL) and crossed-nicols (XPL) photomicrographs of five rock-forming minerals across a full 360-degree rotation, but no classification benchmark has been reported on it. This study establishes that benchmark. Five ImageNet-pre-trained backbones are compared under a rotation-aware, orientation-disjoint evaluation protocol. We find that when many orientations are labelled the five-mineral task saturates at perfect accuracy, so the informative setting is few-shot orientation classification, in which only one or a few labelled rotation views per class are available. In that regime Swin-Tiny is the strongest single-polarisation backbone (96.3 macro-F1), and a minimal two-stream PPL-XPL late-fusion model raises one-shot macro-F1 from 92.2 to 100.0, with the gain concentrated on the orientation-sensitive feldspars, where plagioclase F1 improves by 17.6 points. Feature-level fusion outperforms early channel stacking by 13.8 points. We further verify that the reported accuracy is not an artefact of the splitting protocol, apply Grad-CAM to confirm that predictions attend to diagnostic grain interiors instead of background, and show that the pipeline transfers to a seven-class hand-sample dataset. Code and configurations are released for reproducibility.
Show more
Read Article PDF
Cite
Research Article Open Access
Integrated practice of physical examination and renewal of historical and cultural blocks: a case study of the Shanyin Road Historic Area in Shanghai
Urban physical examination is an important foundational institutional instrument for advancing urban renewal actions and achieving high-quality urban development. At present, urban physical examination practices mainly focus on comprehensive evaluations at the macro district scale, and there is a lack of systematic, specialized physical examination practices for special areas such as historical and cultural conservation areas that entail both conservation requirements and renewal demands. Taking the Shanyin Road Historical and Cultural Conservation Area in Shanghai as the research subject, this paper conducts a specialized renewal physical examination across four dimensions—historical and cultural resources, commercial business vitality, neighborhood spatial quality, and inefficient idle resources. By comprehensively applying multi-source data fusion methods including spatial analysis, field surveys, and questionnaire surveys, it systematically identifies key issues in the conservation area regarding the preservation and repair of historic buildings, optimization of business structures, improvement of neighborhood quality, and activation of idle spaces. Based on the physical examination results, differentiated renewal strategies are proposed for the four main streets—Sichuan North Road, Shanyin Road, Duolun Road, and Tian'ai Road—to explore a refined renewal path of 'physical examination first, renewal second.' The research results can provide practical references for the integrated physical-examination-and-renewal of Shanghai and similar historical conservation areas.
Show more
Read Article PDF
Cite
Research Article Open Access
A statistical framework for reliability evaluation of large language models
Article thumbnail
The growing use of Large Language Models (LLMs) in real-world applications has increased the need for reliable evaluation methods. This study proposes a statistical framework for evaluating LLM reliability by jointly considering model accuracy, hallucination rate, and cross-domain performance stability. Based on the TruthfulQA dataset, 300 test questions were selected using a stratified sampling strategy, and responses were generated by DeepSeek-V4-Flash and GPT-4o-mini. GPT-4o was then employed as an automated evaluator to assess the 600 generated responses. Model reliability was evaluated using three dimensions: Accuracy, Hallucination Rate, and Cross-Domain Stability. In particular, the proposed CDS metric quantifies the consistency of model performance across different knowledge categories. A Weighted Geometric Mean was further employed to construct an overall reliability score. The experimental results show that the proposed framework can identify reliability differences between different LLMs. DeepSeek-V4-Flash achieved an overall reliability score of 0.770, compared with 0.678 for GPT-4o-mini. Sensitivity analysis further shows that the model ranking remains unchanged under different weight configurations, indicating that the proposed framework is reasonably robust to variations in indicator weights. This study provides a statistical framework for quantitatively assessing LLM reliability under multiple evaluation criteria.
Show more
Read Article PDF
Cite