EV Purchase IntentionEV Purchase Intention研究档案RESEARCH NOTEBOOK

消费者问卷 · 经济分析 · 机器学习CONSUMER SURVEY · ECONOMETRICS · MACHINE LEARNING

智能驾驶功能与新能源汽车购买溢价意愿Driver-assistance features and willingness to pay more for an EV

基于 622 份问卷,分析消费者对智能驾驶安全性、功能价值的评价与支付溢价意愿的关系,并比较有序 Logit 和随机森林的预测表现。An analysis of 622 survey responses examining how ratings of driving safety and feature value relate to willingness to pay a premium. Ordered logit and random forest models are also compared using cross-validation.

主规格 primaryPRIMARY SPECIFICATION横截面问卷CROSS-SECTIONAL SURVEY结果版本 2026-09RESULTS · 2026-09
有效样本Respondents—问卷记录survey records
技术认知 ORTechnology OR—主规格 · 控制后primary · controlled
功能价值 ORFunction-value OR—主规格 · 控制后primary · controlled
扩展 QWKExtended QWK—5 折折外预测5-fold held-out prediction
BOOK

研究摘要STUDY SUMMARY

智能驾驶与支付意愿Driver assistance and willingness to pay

622 份问卷中的技术认知、功能价值与支付溢价意愿。Technology perceptions, feature value and willingness to pay a premium in 622 survey responses.

研究档案RESEARCH NOTEBOOK

EV PURCHASE INTENTION / 2026

智能驾驶功能
与购买溢价意愿
Driver-assistance features
and willingness to pay more

消费者问卷 · 研究报告Consumer survey · Results report

01 样本与变量 →01 Sample and variables →p. 01
样本SAMPLE

样本与变量Sample and variables

结果变量是 Q21 的支付溢价意愿。技术认知使用 Q15、Q16,功能价值使用 Q22–Q25;三个指标均保留原问卷的 1–5 分尺度。The outcome is willingness to pay a premium in Q21. Technology recognition uses Q15–Q16 and feature value uses Q22–Q25; all three measures retain the original 1–5 scale.

622份有效问卷survey responses4个结果模块result modules
01 样本与变量 →01 Sample and variables →p. 02
直接关联ASSOCIATIONS

技术认知与功能价值Technology perceptions and feature value

2.08技术认知 T · ORTechnology recognition T · OR
3.83功能价值 V · ORFunction value V · OR

主规格控制性别、年龄、学历、收入、驾龄和驾驶频率;这是横截面关联。The primary specification controls for demographics and driving history; this is a cross-sectional association.

02 直接关联 →02 Direct associations →p. 03
探索性路径INDIRECT PATHS

驾驶乐趣与出行效率Driving pleasure and travel efficiency

分析包括经驾驶乐趣、出行效率以及功能价值的五条间接路径,采用 5,000 次 Bootstrap 估计区间。横截面数据无法确定这些关系的因果顺序。Five indirect paths through driving pleasure, travel efficiency and feature value are estimated with 5,000 bootstrap draws. Cross-sectional data cannot establish the causal ordering of these relationships.

探索性EXPLORATORY
03 间接路径 →03 Indirect paths →p. 04
机器学习MACHINE LEARNING

有序 Logit 与随机森林Ordered logit and random forest

0.685扩展有序 Logit · QWKExtended ordered logit · QWK
0.619随机森林 · QWKRandom forest · QWK

每折测试样本都没有参与该折训练。SHAP 展示特征贡献的大小,不代表因果方向。Each test fold was kept outside training. SHAP shows contribution magnitude, not causal direction.

05 预测与 SHAP →05 Prediction and SHAP →p. 05
研究局限LIMITATIONS

样本与方法的限制Sample and methodological limitations

  • 横截面问卷,不能确定因果。Cross-sectional survey; no causal claim.
  • 比例优势假设尚未专项检验。The proportional-odds assumption needs a dedicated check.
06 方法与来源 →06 Methods and sources →p. 06
01

样本与变量Sample and variables

样本包含 622 份问卷。结果变量为支付溢价意愿,技术认知取 Q15、Q16 的平均值,功能价值取 Q22–Q25 的平均值。The sample contains 622 survey responses. The outcome is willingness to pay a premium. Technology recognition is the mean of Q15 and Q16; feature value is the mean of Q22–Q25.

622原始问卷记录survey records
0审计未发现none in audit
—Q15、Q16Q15 and Q16
—Q22–Q25Q22–Q25
区域题说明:Region item: Q3 区域题是 8 类分类变量,通用 1–5 审计会标记 6–8;Q3 不进入当前主模型。Q3 is an eight-category region item; the generic 1–5 audit flags codes 6–8, and Q3 is not used in the current main model.

Y · 结果变量Y · Outcome

Q21:愿意为智能驾驶功能支付溢价,保留 1–5 的有序等级。Q21: willingness to pay a premium for intelligent-driving functions, retained as an ordered 1–5 outcome.

T · 技术认知T · Technology proxy

主规格为 Q15、Q16 的完整回答均值,代表重要性和安全性认知。The primary proxy is the complete-row mean of Q15 and Q16, representing importance and safety recognition.

V · 功能价值V · Function-value proxy

主规格为 Q22–Q25 的完整回答均值,代表具体辅助驾驶功能的价值认知。The primary proxy is the complete-row mean of Q22–Q25, representing value recognition for concrete assisted-driving functions.

变量构造Variable construction

Y 为 Q21 的支付溢价意愿,保留 1–5 分等级。T 由 Q15“功能重要性”和 Q16“提高安全性”取均值;V 由 Q22–Q25 的自适应巡航、车道保持、自动泊车和拥堵辅助四项取均值。它们分别概括一般技术评价和具体功能评价。M1 为 Q18 的驾驶乐趣评价,M2 为 Q19 的出行效率评价,均为单题指标。Y is the ordered 1–5 response to Q21 on willingness to pay a premium. T averages Q15 (importance) and Q16 (safety). V averages Q22–Q25 on adaptive cruise control, lane keeping, parking assistance and traffic-jam assistance. M1 is Q18 (driving pleasure); M2 is Q19 (travel efficiency), both single-item measures.

样本处理与测量含义Sample preparation and measurement

从原始题项读取数据,不使用预先生成的处理列。组合指标要求所用题项回答完整;各分析按所需变量的完整记录估计。性别、年龄、学历、收入、驾龄和驾驶频率按类别编码。T 和 V 的 Cronbach α 分别为 0.858、0.871,反映题项内部一致性,但不能单独证明测量效度。V 的原题涉及功能对购买意愿的影响,因此与 Y 的含义有接近之处。The pipeline reads original items rather than pre-generated processed columns. Composites require complete item responses, and each analysis uses complete records for its required variables. Six demographic and driving controls are category-coded. Cronbach’s alpha is 0.858 for T and 0.871 for V; these assess internal consistency, not construct validity. V’s questions already refer to purchase intention, so its meaning overlaps with Y.

02

直接关联Direct associations

主模型使用有序 Logit,并控制性别、年龄、学历、收入、驾龄和驾驶频率。OR 大于 1 表示更高的 T/V 与更高购买溢价意愿的排序相关。The primary ordered-logit model controls for gender, age, education, income, driving experience, and driving frequency. An OR above 1 indicates an association with a higher willingness-to-pay category.

模型设定Model specification

研究问题是:在控制人口特征与驾驶经历后,技术认知和功能价值是否仍与支付溢价意愿有关?Y 有五个有序等级,因此使用有序 Logit,同时纳入 T、V 和六类控制变量。T、V 保持原来的 1–5 分尺度,不做标准化;表中优势比对应指标增加 1 分。The question is whether T and V remain associated with willingness to pay after accounting for demographics and driving history. Ordered logit models the five ordered Y categories using T, V and six categorical controls. T and V remain on their original 1–5 scales; odds ratios refer to a one-point increase.

主规格优势比Primary odds ratios

点估计来自 primary / controlled;虚线为 OR=1。置信区间见下表。Estimates are from primary / controlled; the dashed line is OR=1. Confidence intervals are shown below.

主模型估计Primary model estimates

变量VariableOR95% CIpN

优势比表示控制其他变量后的关联,因果方向无法由本次问卷确定。The odds ratios describe adjusted associations. The survey does not establish their causal direction.

估计结果Estimated associations

在 622 份记录中,T 的 OR 为 2.08(95% CI:1.62–2.68),V 为 3.83(2.84–5.16),两项 p 值均小于 0.001。按模型的比例优势假设,其他变量不变时,T 增加 1 分对应跨越任一 Y 等级阈值的优势乘以约 2.08。OR 不是概率倍数,也不是支付金额的增幅。V 的点估计更大,但其题项与 Y 更接近,不能仅凭 OR 排定真实影响大小。Across 622 records, T has OR 2.08 (95% CI 1.62–2.68) and V has OR 3.83 (2.84–5.16), both p<0.001. Under proportional odds, a one-point increase in T multiplies the odds of exceeding any Y threshold by about 2.08, holding other variables fixed. This is neither a probability ratio nor a change in payment amount. V’s larger estimate must also be read in light of its conceptual overlap with Y.

03

间接路径Indirect paths

五条间接路径均使用 5,000 次 Bootstrap 估计 95% 置信区间。结果反映本样本中的间接关联;因果中介关系需要额外的研究设计和假设支持。The five indirect paths use 5,000 bootstrap draws to estimate 95% confidence intervals. The estimates describe indirect associations; causal mediation would require additional design and identification assumptions.

路径与计算方法Paths and estimation

分别估计 T→M1→Y、T→M2→Y、V→M1→Y、V→M2→Y 和 T→V→Y。每条路径单独拟合:先用 X 和控制变量解释中间变量 M,得到 a;再用 X、M 和控制变量解释 Y,得到 b;间接关联取 a×b。两个方程均为含截距的 OLS,把评分作为连续变量近似。按受访者有放回抽样 5,000 次,使用乘积估计分布的 2.5% 和 97.5% 分位数构造区间。Five paths are fitted separately: T→M1→Y, T→M2→Y, V→M1→Y, V→M2→Y and T→V→Y. For each, OLS regresses M on X and controls to estimate a, and Y on X, M and controls to estimate b. Both equations include intercepts and approximate ratings as continuous. The product a×b is bootstrapped by resampling respondents with replacement 5,000 times; percentile intervals use the 2.5th and 97.5th percentiles.

间接关联与区间Indirect associations and intervals

点为间接关联估计,横线为 95% Bootstrap 区间;每条完整路径均以 Y 为终点。Points show indirect estimates; horizontal lines show 95% bootstrap intervals. Every full path ends at Y.

Bootstrap 路径表Bootstrap path table

路径Pathab95% CI有效/总数Valid / total

中介方程是探索性的 OLS 近似,不能与有序 Logit 系数直接相乘。Mediator equations are exploratory OLS approximations and are not multiplied by ordered-logit coefficients.

路径结果Path estimates

五条路径的区间均高于 0。经驾驶乐趣的估计分别为 0.256 和 0.257,经出行效率为 0.180 和 0.163;T→V→Y 为 0.351。每条路径均得到 5,000 次有效估计。这些数值属于独立的线性路径模型,不与前节的 OR 相加或相乘,也不能加总为“解释比例”。同一时点的问卷无法确认变量发生顺序,因此这里报告探索性间接关联。All five intervals lie above zero. Products through driving pleasure are 0.256 and 0.257, through travel efficiency 0.180 and 0.163, and through T→V→Y 0.351. Each path has 5,000 valid bootstrap estimates. These are separate linear path models; their estimates cannot be combined with the ordered-logit ORs or summed into an explained proportion. The cross-sectional survey cannot establish temporal ordering.

04

人群差异Group differences

性别、年龄、收入、驾龄和驾驶频率分别进行嵌套有序 Logit LR 检验。五项检验的 p 值使用 Holm 方法进行多重比较校正。Gender, age, income, driving experience, and driving frequency are tested with nested ordered-logit LR comparisons. The five p-values are adjusted for multiple comparisons using the Holm method.

分组检验Group comparisons

检验技术认知、功能价值与 Y 的关系是否随性别、年龄、收入、驾龄或驾驶频率而变化。每次比较一个含分组主项的有序 Logit 与加入 T×分组、V×分组交互项的模型,使用似然比(LR)检验联合判断斜率差异。五个分组分别检验,再用 Holm 方法校正多重比较。表中 k 为该次分析的组数。For each grouping, a restricted ordered logit containing group main effects is compared with a model adding T-by-group and V-by-group interactions. The likelihood-ratio test jointly assesses slope differences. Five grouping tests are adjusted using Holm’s method. The table’s k is the number of groups.

Holm 校正后的 p 值Holm-adjusted p-values

虚线为 0.05 判断线。本次五项比较校正后均未低于该线。The dashed line marks 0.05. None of the five adjusted comparisons falls below it.

异质性检验表Heterogeneity table

分组GroupingkLRpHolm p

五项检验经 Holm 校正后均未达到 0.05 显著性水平。None of the five tests meets the 0.05 significance threshold after Holm adjustment.

分组结果Group results

性别和年龄的原始 p 值分别为 0.013、0.042,校正后为 0.063、0.170;其余分组的校正值也高于 0.05。因此,本次分析没有得到经多重比较校正后仍显著的人群差异证据。这不等于证明各组完全相同,也不支持依据个别未校正结果给出针对某个人群的结论。Raw p-values for gender and age are 0.013 and 0.042, increasing to 0.063 and 0.170 after adjustment. All other adjusted values also exceed 0.05. The analysis therefore finds no group differences meeting the adjusted significance threshold. This does not establish that groups are identical.

05

预测与 SHAPPrediction and SHAP

五折折外评估显示,加入 T、V、M1、M2 后模型明显超过多数类基线;在本次运行中,有序 Logit 的 QWK 高于随机森林。Five-fold held-out evaluation shows a clear improvement over the majority baseline after adding T, V, M1, and M2. In this run, ordered logit has a higher QWK than the random forest.

预测实验Prediction experiment

预测目标仍为 Y 的五个等级。比较三组输入:仅控制变量、控制变量+T/V、控制变量+T/V/M1/M2;每组均比较多数类基线、有序 Logit 和随机森林。使用随机种子 42 的五折分层交叉验证,每折约四份数据训练、一份评价。随机森林固定为 200 棵树、最大深度 6;表中报告五折指标的平均值。The target remains the five Y categories. Three feature sets—controls, controls plus T/V, and controls plus T/V/M1/M2—are evaluated with a majority baseline, ordered logit and random forest. Five-fold stratified cross-validation uses seed 42, training on four folds and evaluating on the remaining fold. Random forests use 200 trees and maximum depth 6. The table reports fold means.

折外 QWKHeld-out QWK

多数类基线 QWK 为 0;扩展有序 Logit 约为 0.685。The majority baseline has QWK 0; extended ordered logit is about 0.685.

模型比较Model comparison

特征集Feature set模型ModelAcc.F1QWKMAE

SHAP 特征排序SHAP feature ranking

图中展示扩展随机森林对 P(Y≥4) 的全部 23 个输入特征,按平均绝对 SHAP 排序。数值表示贡献大小,不表示作用方向或因果关系。The chart shows all 23 inputs used by the extended forest for P(Y≥4), ranked by mean absolute SHAP. Magnitude does not indicate direction or causality.

交叉验证:Cross-validation: 加入全部研究变量后,有序 Logit 的平均准确率为 54.5%,随机森林为 53.5%;平均 QWK 分别为 0.685 和 0.619。每一折测试样本都没有参与该折训练。With all study variables included, mean accuracy is 54.5% for ordered logit and 53.5% for random forest; mean QWK is 0.685 and 0.619. Each test fold was kept outside training.

指标与结果Metrics and results

Accuracy 统计等级完全预测正确的比例;F1 为各等级 F1 的宏平均;QWK 为二次加权 Kappa,考虑等级误差并校正随机一致性;MAE 为平均相差几个等级。仅控制变量的预测接近基线。加入 T/V 后,有序 Logit 的 QWK 为 0.650;再加入 M1/M2 后为 0.685,准确率为 54.5%。同样扩展输入下,随机森林为 0.619 和 53.5%,本次结果没有显示它优于有序 Logit。Accuracy measures exact category matches; F1 is macro-averaged across categories; QWK is quadratic-weighted kappa; MAE is the mean absolute category error. Controls alone perform near the baseline. Ordered-logit QWK reaches 0.650 with T/V and 0.685 with M1/M2 added, with accuracy 54.5%. The extended random forest scores 0.619 and 53.5%, so it does not outperform ordered logit in this run.

SHAP 的解释对象What SHAP explains

对扩展随机森林预测“Y≥4”的概率计算 SHAP,在每折未参与训练的样本上解释,再按平均绝对值汇总排序。功能价值、技术认知、驾驶乐趣和出行效率排在前四。图中也列出分类控制变量编码后的各项,所以共有 23 个输入特征。排序反映模型预测依赖程度,不表示因果效应;交叉验证将每折训练样本与测试样本分开。SHAP explains the extended forest’s predicted probability of Y≥4 on each held-out fold. Features are ranked by mean absolute attribution, with V, T, M1 and M2 first. Category-coded controls bring the total to 23 inputs. The ranking describes predictive attribution, not causal effects. Each fold keeps its test sample outside training.

06

方法与来源Methods and sources

数据来自 622 份消费者问卷。完整分析代码与双语 Notebook 见项目仓库。The study uses 622 consumer survey responses. Analysis code and bilingual notebooks are available in the repository.

01

变量定义Variable definitions

题项、组合变量、控制变量和缺失值处理规则见 src/data_schema.py 与 src/config.py。Item mappings, composites, controls and missing-value rules are defined in src/data_schema.py and src/config.py.

02

主规格Primary model

正文统一使用当前主规格:T 由 Q15、Q16 构成,V 由 Q22–Q25 构成,控制变量按类别编码。The report uses the current primary specification: T uses Q15 and Q16, V uses Q22–Q25, and controls are category-coded.

03

折外验证Held-out checks

有序 Logit、随机森林和多数类基线使用同一组分层折。Ordered logit, random forest, and the majority baseline share the same stratified folds.

04

研究范围Study scope

结果反映这份横截面问卷中的关联和预测表现,不能确定因果关系。Results describe associations and predictive performance in this cross-sectional survey. They do not establish causality.