The Science and Information (SAI) Organization publishes open-access peer-reviewed journals in computer science and artificial intelligence.

Contact Info
Website thesai.org
Follow Us
Contact Info
Follow Us
Research Article | Open Access |
First page preview

Characterizing Operational Drift in Cement Manufacturing Process Data via Self-Supervised Representation Learning

Author 1: Changgyun Kim
International Journal of Advanced Computer Science and Applications (IJACSA) · Vol. 17, No. 6 · Published 2026

DOI: https://doi.org/10.14569/IJACSA.2026.0170607

Abstract

Cement finish milling generates large volumes of process variable data at hourly resolution, while quality measurements (Blaine fineness, 44 µm residue) are recorded only every two to four hours, producing a sparse-label regime with substantial unlabeled data accumulated over multi-year operation. In this study, we analyze 188,858 hourly records collected from three industrial cement mill units spanning 2017 to 2025 (approximately nine years) and characterize the operational drift structure embedded in the data. Using Pruned Exact Linear Time (PELT) change-point detection on three critical process variables, we identify between 12 and 21 detected drift events per mill, with Cohen's d effect sizes exceeding 1.0 for the majority of detected change points and reaching 4.3 in extreme cases. We then apply SCARF, a contrastive self-supervised learning method for tabular data, to learn 128-dimensional representations on the combined labeled (51,225 records) and unlabeled (111,506 records) data. Multi-seed training yields stable validation InfoNCE loss of 6.93 ± 0.02. Three clustering algorithms applied to the learned embedding space (K-means, Gaussian Mixture Model, HDBSCAN) consistently select 15 to 21 operational regimes, with silhouette scores between 0.36 and 0.41. The Adjusted Rand Index between embedding-space K-means and process variable space K-means is 0.35, indicating that the learned representation preserves coarse regime structure while resolving finer sub-regime variability. Cluster analysis further reveals strong mill specificity, with 14 of 15 embedding clusters dominated by a single mill, and temporal cluster evolution that aligns with PELT-detected change-point boundaries. These findings establish that long-term cement process data contains a richer operational regime structure than implied by raw process variable clustering, and that self-supervised pretraining can recover this structure, and that the resulting representation yields statistically significant gains in downstream quality prediction.

Keywords

How to Cite this Article

Changgyun Kim. "Characterizing Operational Drift in Cement Manufacturing Process Data via Self-Supervised Representation Learning". International Journal of Advanced Computer Science and Applications (IJACSA), Vol. 17, No. 6, 2026. https://doi.org/10.14569/IJACSA.2026.0170607

BibTeX

@article{Kim2026,
  title     = {Characterizing Operational Drift in Cement Manufacturing Process Data via Self-Supervised Representation Learning},
  journal   = {International Journal of Advanced Computer Science and Applications},
  volume    = {17},
  number    = {6},
  year      = {2026},
  publisher = {The Science and Information Organization},
  author    = {Changgyun Kim},
  doi       = {10.14569/IJACSA.2026.0170607},
  url       = {https://doi.org/10.14569/IJACSA.2026.0170607}
}

Open Access — licensed under a Creative Commons Attribution 4.0 International License. Unrestricted use, distribution, and reproduction in any medium, even commercially, as long as the original work is properly cited.