Grouping hundreds of cryptocurrencies by behavior manually leads to confusion — assets with similar volatility may have different correlation with BTC, and trading volumes change unpredictably. Our clustering model leverages cryptocurrency correlation and volatility analysis to identify hidden groups. Our engineers, with 5+ years of market experience, 7 years of ML experience and 50+ completed crypto market analysis projects, automate this process. The model identifies hidden groups based on 15+ metrics: annualized return, volatility, Sharpe ratio, skewness, kurtosis, VaR 95%, CVaR 95%, maximum drawdown, correlation with BTC, 30-day momentum, average daily volume. This is not classification — we don't assign labels, we find natural groups by combining assets with similar behavioral patterns.
Such clustering allows portfolio diversification by picking 1-2 assets from each cluster, reducing correlation. It also supports rotational strategies: when one asset in a cluster surges, we look for lagging assets in the same cluster. Finally, it helps understand market structure, identifying groups like "blue chips", "high-beta altcoins", and "decorrelated assets".
How We Build Features for Clustering
Feature engineering is the key step. We take hourly close prices for the last 90 days. For each asset, we calculate:
- annualized return and volatility,
- Sharpe ratio (return to risk ratio),
- 30-day momentum,
- maximum drawdown,
- correlation with BTC (if available),
- average daily volume in USD.
Code for creating features:
import pandas as pd
import numpy as np
from sklearn.preprocessing import StandardScaler
def create_behavioral_features(prices_dict, lookback_days=90):
features = {}
for symbol, price_series in prices_dict.items():
returns = price_series.pct_change().dropna()
if len(returns) < lookback_days * 24: # hourly data
continue
recent_returns = returns.iloc[-lookback_days*24:]
features[symbol] = {
# Return characteristics
'annualized_return': recent_returns.mean() * 365 * 24,
'annualized_vol': recent_returns.std() * np.sqrt(365 * 24),
'sharpe': recent_returns.mean() / (recent_returns.std() + 1e-8) * np.sqrt(365*24),
# Distribution shape
'skewness': recent_returns.skew(),
'kurtosis': recent_returns.kurt(),
# Tail risk
'var_95': np.percentile(recent_returns, 5),
'cvar_95': recent_returns[recent_returns <= np.percentile(recent_returns, 5)].mean(),
# Trend characteristics
'momentum_30d': price_series.iloc[-720:].pct_change(720).iloc[-1], # 30d return
'trend_strength': abs(recent_returns.mean()) / (recent_returns.std() + 1e-8),
# Drawdown
'max_drawdown': calculate_max_drawdown(price_series.iloc[-lookback_days*24:]),
# Correlation with BTC (if available)
'btc_corr': recent_returns.corr(prices_dict.get('BTC', pd.Series()).pct_change().dropna()),
# Volume-based (if volume data available)
'avg_daily_volume_usd': get_avg_daily_volume(symbol),
}
return pd.DataFrame(features).T
In one project for a hedge fund, we clustered 150 cryptocurrencies in 2 weeks for a cost of $5,000. As a result, the client built a portfolio with a 0.35 correlation between assets from different clusters, reducing drawdown risk by 40%. This saved over 200 hours of manual analysis annually, delivering a 10x ROI.
Key Metrics for Clustering
Not all metrics are equally useful. Correlation with BTC and volatility often dominate, but adding momentum and drawdown improves separation of speculative assets. We apply PCA for feature importance analysis and remove multicollinear features.
Comparison of Clustering Algorithms
We use three approaches — each with its strengths.
K-Means — a classic: fast (3x faster than DBSCAN on 500+ objects), but assumes spherical clusters of equal size. Suitable for initial partitioning.
from sklearn.cluster import KMeans
from sklearn.preprocessing import StandardScaler
def kmeans_clustering(features_df, n_clusters=6, seed=42):
scaler = StandardScaler()
features_scaled = scaler.fit_transform(features_df.fillna(0))
inertias = []
k_range = range(2, 15)
for k in k_range:
km = KMeans(n_clusters=k, random_state=seed, n_init=10)
km.fit(features_scaled)
inertias.append(km.inertia_)
best_k = find_elbow(inertias, k_range)
km = KMeans(n_clusters=best_k, random_state=seed, n_init=10)
labels = km.fit_predict(features_scaled)
return labels, km, scaler
DBSCAN — does not require specifying the number of clusters, detects outliers (noise points). Good when clusters have complex shapes.
from sklearn.cluster import DBSCAN
def dbscan_clustering(features_scaled, eps=0.5, min_samples=3):
db = DBSCAN(eps=eps, min_samples=min_samples, metric='euclidean')
labels = db.fit_predict(features_scaled)
n_clusters = len(set(labels)) - (1 if -1 in labels else 0)
n_noise = (labels == -1).sum()
return labels, n_clusters, n_noise
Hierarchical clustering — builds a dendrogram, visually showing hierarchy. We use it for visual analysis when nested clusters need to be seen.
UMAP is Better Than PCA for Cluster Visualization
To visualize high-dimensional data, we reduce dimensionality to 2D. UMAP (Uniform Manifold Approximation and Projection), unlike linear PCA, better preserves global and local structure. In practice, UMAP yields more compact and separated clusters, especially for data with nonlinear dependencies. UMAP is described in the work by McInnes et al. (see Wikipedia).
from sklearn.decomposition import PCA
from sklearn.manifold import TSNE
import umap
def reduce_dimensions(features_scaled, method='umap', n_components=2):
if method == 'pca':
reducer = PCA(n_components=n_components, random_state=42)
elif method == 'tsne':
reducer = TSNE(n_components=n_components, random_state=42,
perplexity=min(30, len(features_scaled)//4))
elif method == 'umap':
reducer = umap.UMAP(n_components=n_components, random_state=42,
n_neighbors=min(15, len(features_scaled)//3))
embedding = reducer.fit_transform(features_scaled)
return embedding
After reduction, we create a scatter plot with color coding by cluster — the main visualization of results.
Interpreting Clusters
After clustering, we analyze the mean values of metrics per cluster. For example:
| Cluster | Correlation with BTC | Volatility (annualized) | Sharpe Ratio | Interpretation |
|---|---|---|---|---|
| 0 | >0.85 | >1.5 | <1.0 | High-beta altcoins |
| 1 | >0.8 | <1.0 | >1.5 | Blue-chip crypto |
| 2 | <0.5 | <0.8 | >2.0 | Decorrelated assets |
| 3 | <0.5 | >2.0 | <0.5 | Speculative / memes |
Code for automatic interpretation:
def describe_clusters(features_df, labels):
features_df['cluster'] = labels
cluster_stats = features_df.groupby('cluster').agg({
'annualized_return': 'mean',
'annualized_vol': 'mean',
'sharpe': 'mean',
'btc_corr': 'mean',
'max_drawdown': 'mean',
'skewness': 'mean'
}).round(3)
cluster_names = {}
for cluster_id, row in cluster_stats.iterrows():
if row['btc_corr'] > 0.85 and row['annualized_vol'] > 1.5:
name = 'High-beta altcoins'
elif row['btc_corr'] > 0.8 and row['annualized_vol'] < 1.0:
name = 'Blue-chip crypto'
elif row['btc_corr'] < 0.5:
name = 'Decorrelated assets'
elif row['sharpe'] > 2.0:
name = 'Strong performers'
else:
name = f'Cluster {cluster_id}'
cluster_names[cluster_id] = name
return cluster_stats, cluster_names
Comparison of clustering algorithms
| Algorithm | Speed | Requires k | Outlier detection | Cluster shape |
|---|---|---|---|---|
| K-Means | Fast | Yes (k) | No | Spherical |
| DBSCAN | Medium | No (eps, min_samples) | Yes | Arbitrary |
| Hierarchical | Slow | No (k) | No | Any (dendrogram) |
Work Process
- Data collection and cleaning — historical prices, volumes, on-chain metrics for 90-180 days.
- Feature engineering — calculate 15+ metrics, normalization, feature selection.
- Algorithm selection — test K-Means, DBSCAN, Hierarchical; optimize number of clusters.
- Validation — silhouette score, UMAP visualization, cluster stability.
- Visualization — dashboard with interactive cluster map.
- Documentation — interpretation of each cluster, update instructions.
What's Included
- Python model code with comments.
- Dashboard (Plotly/Dash) with cluster visualization.
- Documentation interpreting each cluster.
- Model update guide (recommended monthly updates).
- 1 month of support after delivery.
Leveraging our 7 years of ML experience and 50+ completed projects, we guarantee high-quality work. Typical development cost ranges from $3,000 to $7,000 depending on the number of assets and data depth; we determine it after analyzing your dataset. Contact us for a preliminary assessment — we'll design an architecture tailored to your task. Order the development of a clustering model and get a ready-to-use tool for portfolio diversification.







