线性度最小二乘法线性计算

Feature selection is a process where the predictor variables that contribute most significantly towards the prediction/ classification of the target variable are selected. In feature selection for linear regression models, we are concerned with four aspects regarding the variables. Framed as a mnemonic “LINE”, these are:

特征选择是一个过程,其中选择对目标变量的预测/分类贡献最大的预测变量。 在线性回归模型的特征选择中,我们关注与变量有关的四个方面。 这些框架以助记符“ LINE”表示:

  1. Linearity. The selected variable possesses a linear relationship with the target variable.

    大号 inearity。 所选变量与目标变量具有线性关系。

  2. Independence of predictor variables. Selected variables to be independent of each other.

    独立于预测变量。 选择的变量要彼此独立。

  3. Normality. Residuals generally follow a normal distribution (mean of zero).

    ñormality。 残差通常遵循正态分布(均值为零)。

  4. Equality of variance. The residual errors are generally consistent across the values of predictor variables (i.e. Homoscedasticity).

    E方差质量。 残余误差通常在预测变量的值之间是一致的(即同方差)。

In cases where selected predictor variables are not independent of each other, we would not be able to clearly determine or attribute the contribution from the various predictor variables towards the target variable — interpretability of the model coefficients becomes an issue.

在选定的预测变量彼​​此不独立的情况下,我们将无法清晰地确定或将各种预测变量对目标变量的贡献归因于模型系数的可解释性。

One approach in feature selection would be through the use of p-values; where variables with p-values above a certain threshold (typically +/- 0.05) are surfaced as not significantly contributing towards the target variable, and hence can be dropped to reduce model complexity. However, this approach has its own challenges when multicollinearity exists among the predictor variables as illustrated through the Boston housing dataset in sklearn.

特征选择的一种方法是使用p值。 其中p值高于某个阈值(通常为+/- 0.05)的变量会因为对目标变量的贡献不大而浮出水面,因此可以将其删除以降低模型的复杂性。 但是,当预测变量之间存在多重共线性时,如sklearn中的Boston住房数据集所示,这种方法面临着自己的挑战。

# Original (full) set of variables
X_o = df_wdummy[['CRIM', 'ZN', 'INDUS', 'NOX', 'RM', 'AGE', 'DIS', 'RAD', 'TAX', 'PTRATIO', 'B', 'LSTAT', 'CHAS_1.0']]
X_o = sm.add_constant(X_o)
y_o = df_wdummy['MEDV']
# Baseline results
model_o = sm.OLS(y_o, X_o)
results_o = model_o.fit()
results_o.summary()
Image for post
OLS regression results 1, note the p-values above 0.05 (image by author)
OLS回归结果1,请注意p值大于0.05(作者提供)

Based on the p-values, we remove variables INDUS and AGE and review the updated p-values.

基于p值,我们删除变量INDUS和AGE并查看更新的p值。

# Remove Age -> Remove INDUS, based on p-values.
X_r0 = df_wdummy[['CRIM', 'ZN', 'NOX', 'RM', 'DIS', 'RAD', 'TAX',
'PTRATIO', 'B', 'LSTAT', 'CHAS_1.0']]
X_r0 = sm.add_constant(X_r0)
y_r0 = df_wdummy['MEDV']
# results
model_r0 = sm.OLS(y_r0, X_r0)
results_r0 = model_r0.fit()
results_r0.summary()
Image for post
Updated OLS regression results; note none of the p-values above 0.05 (image by author)
更新了OLS回归结果; 请注意,p值均不得高于0.05(作者提供的图片)

Though all p-values are lower than 0.05, multicollinearity still exists (between TAX and RAD variables.) The presence of multicollinearity can mask the importance of the respective variable contributions to the target variable, where the interpretability of p-values then becomes challenging. We could use correlation measures and matrices to help visualize and mitigate multicollinearity. Such an approach is fine until we need to use different correlation measures (i.e. Spearman, Pearson, Kendall) due to the inherent attributes of the variables. In the example above, the variable RAD (index of accessibility to radial highways) is an ordinal variable. TAX (full-value property-tax rate per $10,000) is a continuous variable (not normally distributed). Using the different correlation measures and matrices, one could potentially overlook correlation among different categories of variables.

尽管所有p值均低于0.05,但多重共线性仍然存在(在TAX和RAD变量之间) 。多重共线性的存在会掩盖各个变量对目标变量的重要性,因此 p值的可解释性将具有挑战性。 我们可以使用相关度量和矩阵来帮助可视化和减轻多重共线性。 这种方法很好,直到由于变量的固有属性而需要使用不同的相关度量(即Spearman,Pearson,Kendall)。 在上面的示例中,变量RAD(径向公路的可达性指数)是一个序数变量。 TAX(每10,000美元的全值财产税率)是一个连续变量(非正态分布)。 使用不同的相关度量和矩阵,可能会忽略不同类别变量之间的相关性。

Another approach to identify multicollinearity is via the Variance Inflation Factor. VIF indicates the percentage of the variance inflated for each variable’s coefficient. Beginning at a value of 1 (no collinearity), a VIF between 1–5 indicates moderate collinearity while values above 5 indicate high collinearity. Some cases where high VIF would be acceptable include the use of interaction terms, polynomial terms, or dummy variables (nominal variables with three or more categories). Correlation matrices enable identification of correlation among variable pairs while VIF enables the overall assessment of multicollinearity. The correlation matrix for most of the continuous variables is presented below to highlight the various collinear variable pairs. VIF can be calculated using the statsmodels package; the code block below presents the VIF values with collinear variables included (left) and removed (right).

识别多重共线性的另一种方法是通过方差膨胀因子。 VIF 表示每个变量系数的方差膨胀百分比。 从值1(无共线性)开始,VIF在1-5之间表示中等共线性,而值大于5则表示高共线性。 高VIF可以接受的一些情况包括使用交互项,多项式项或伪变量(具有三个或更多类别的名义变量)。 相关矩阵可以识别变量对之间的相关性,而VIF可以对多重共线性进行整体评估。 下面列出了大多数连续变量的相关矩阵,以突出显示各种共线变量对。 可以使用statsmodels包来计算VIF; 下面的代码块显示了带有共线变量的VIF值(左)和被删除(右)。

Image for post
correlation matrix for the continuous variables; Kendall coefficient (image by author)
连续变量的相关矩阵; 肯德尔系数(作者提供)
# Setting the predictor variables
X_o = df_wdummy[['CRIM', 'ZN', 'INDUS', 'NOX', 'RM', 'AGE', 'DIS', 'RAD', 'TAX', 'PTRATIO', 'B', 'LSTAT', 'CHAS_1.0']]
X_r1 = df_wdummy[['CRIM', 'ZN', 'INDUS', 'RM', 'AGE', 'DIS', 'TAX', 'PTRATIO', 'B', 'LSTAT', 'CHAS_1.0']]#
from statsmodels.stats.outliers_influence import variance_inflation_factorvif = pd.Series([variance_inflation_factor(X_o.values, i) for i in range(X_o.shape[1])], index=X_o.columns,
name='vif_full')
vif_r = pd.Series([variance_inflation_factor(X_r1.values, i) for i in range(X_r1.shape[1])], index=X_r1.columns,
name='vif_collinear_rvmd')
pd.concat([vif, vif_r], axis=1)
Image for post
VIF values for predictor variables (image by author)
预测变量的VIF值(作者提供的图像)

The VIF values correspond with the correlation matrix; for example, variable-pair NOX and INDUS, the correlation coefficient is above 0.5 (0.61), and the respective VIF values are above 5. The removal of the collinear variables RAD and NOX improved the VIF figures. Dropping collinear variables solely based on the highest VIF figure is not a guaranteed way towards building the best performing model, as elaborated in the next section.

VIF值与相关矩阵相对应; 例如,变量对NOX和INDUS,相关系数大于0.5(0.61),相应的VIF值大于5。共线变量RAD和NOX的删除改善了VIF值。 仅根据最高VIF值来丢弃共线性变量并不是建立最佳性能模型的保证方法,这将在下一节中详细说明。

We build a baseline model by dropping all collinear variables identified in the correlation matrix (shown above), pending TAX and RAD to be dropped next.

我们通过删除相关矩阵中标识的所有共线变量(如上所示)来构建基线模型,然后再删除未决的TAX和RAD。

# Baseline variables 
X_bl = df_wdummy[['INDUS', 'RM', 'AGE', 'RAD', 'TAX', 'PTRATIO', 'LSTAT']]
y_bl = df_wdummy['MEDV']# Explore mitigating multi-collinearity
vif_bl = pd.Series([variance_inflation_factor(X_bl.values, i) for i in range(X_bl.shape[1])], index=X_bl.columns,
name='vif_bl')X_noTAX = X_bl.drop(['TAX'],axis=1)
X_noRAD = X_bl.drop(['RAD'],axis=1)
vif_noTAX = pd.Series([variance_inflation_factor(X_noTAX.values, i) for i in range(X_noTAX.shape[1])],
index=X_noTAX.columns, name='vif_noTAX')
vif_noRAD = pd.Series([variance_inflation_factor(X_noRAD.values, i) for i in range(X_noRAD.shape[1])],
index=X_noRAD.columns, name='vif_noRAD')
pd.concat([vif_bl, vif_noTAX, vif_noRAD], axis=1)
Image for post
Assessing VIF figures of predictor variables in baseline model (image by author)
评估基线模型中预测变量的VIF数字(作者提供的图像)

While it appears dropping the TAX variable based on VIF seems to be better, a prudent approach is to check via the adjusted R-squared metric (adjusted for the number of predictors, the metric increases only if the next added variable improves the model more than would be expected by chance).

虽然似乎删除基于VIF的TAX变量似乎更好,但一种审慎的方法是通过调整后的R平方度量(针对预测变量的数量进行检查,仅当下一个添加的变量对模型的改进大于会是偶然的)。

# Without TAX
model = sm.OLS(y, sm.add_constant(X_noTAX)).fit()
print_model = model.summary()
print(print_model)
Image for post
Summary of model metrics (w/o TAX) (image by author)
模型指标摘要(不含TAX)(作者提供图片)
# Without RAD
model = sm.OLS(y, sm.add_constant(X_noRAD)).fit()
print_model = model.summary()
print(print_model)
Image for post
Summary of model metrics (w/o RAD) (image by author)
模型指标摘要(不含RAD)(作者提供图片)

From the higher adjusted R-squared figure, we can infer that the model performs better with the RAD variable dropped! With multicollinearity issue addressed, the next step could be to explore the addition of interaction terms to potentially boost model performance.

从较高的调整后R平方图,我们可以推断出RAD变量删除后该模型的性能更好! 解决了多重共线性问题后,下一步可能是探索增加交互项以潜在地提高模型性能。

In summary, the presence of multicollinearity can mask the importance of predictor variables to the target variable. The use of both Correlation matrices and VIF can help identify correlated variable pairs and assess multicollinearity among selected variables (features). While some iterations would still be necessary to assess model performance, with VIF and correlation matrices, we would be able to make a better-informed decision for feature selection.

总之,多重共线性的存在可以掩盖预测变量对目标变量的重要性。 相关矩阵和VIF的使用可以帮助识别相关变量对并评估所选变量(特征)之间的多重共线性。 尽管使用VIF和相关矩阵仍需要一些迭代来评估模型性能,但我们将能够为功能选择做出更明智的决策。

Codes are hosted here: https://github.com/AngShengJun/petProj/tree/master/eda_viz

代码托管在这里: https : //github.com/AngShengJun/petProj/tree/master/eda_viz

翻译自: https://towardsdatascience.com/collinearity-measures-6543d8597a2e

线性度最小二乘法线性计算

Logo

DAMO开发者矩阵,由阿里巴巴达摩院和中国互联网协会联合发起,致力于探讨最前沿的技术趋势与应用成果,搭建高质量的交流与分享平台,推动技术创新与产业应用链接,围绕“人工智能与新型计算”构建开放共享的开发者生态。

更多推荐