نوع مقاله : مقاله پژوهشی
عنوان مقاله English
نویسندگان English
Accurate customer knowledge and robust segmentation are prerequisites for data-driven banking. This study designs and evaluates a scalable clustering model for individual bank customers using extended RFM concepts and MiniBatchKMeans. Operational data from a private bank were aggregated at customer level and processed through data-quality controls, signed logarithmic transformations for skewed financial variables, outlier control, and standardized preprocessing fitted without leakage. The final clustering space contained ten robust and traceable features: signed-log financial value, profitability and average balance; log-transformed recency and frequency; digital ratio and digital intensity; channel diversity and service diversity; and a robust profit-margin indicator. Because only 257 positive risk events were observed, risk variables were excluded from the clustering space and retained as a separate post-clustering governance layer. Cluster counts from k=2 to k=8 were assessed using Silhouette, Davies-Bouldin, Calinski-Harabasz, multi-seed ARI, aligned Jaccard stability, and minimum cluster size. The final two-cluster solution achieved Silhouette=0.5519, Davies-Bouldin=0.9009, Calinski-Harabasz=11999.13, mean ARI=0.9976 and mean Jaccard=0.9986. Among 378,312 clusterable customers, 66.74% belonged to a less-digital/lower-financial-value group and 33.26% to a more-digital/higher-financial-value group; 9,307 records with simultaneously zero financial base were kept in a separate data-quality queue rather than interpreted as a customer segment. The results support MiniBatchKMeans as a scalable and interpretable tool for banking customer segmentation, while emphasizing that cluster labels are descriptive and should not be treated as causal, credit-risk, or future-behavior predictions.
کلیدواژهها English