Applied Statistics and Experimental Design
Time Series Stock Price Analysis
A comprehensive statistical analysis and forecasting framework
for time series stock price data
Class/Group: Group 23
Program: Data Science and Artificial Intelligence (DS-AI)
Team Members:
1. Nguyen Huu Duc Anh
2. To Le Quang
3. Nguyen Manh Hung
4. Nguyen Duy Long
5. Nguyen Ngoc Tuan Anh
Hanoi, 2026
StockPriceAnalysis 1
Contents
1 Introduction 3
1.1 Abstract . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3
1.2 RelatedWorks. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4
1.2.1 ClassicalStatisticalApproaches . . . . . . . . . . . . . . . . . . . . 4
1.2.2 MachineLearninginFinance . . . . . . . . . . . . . . . . . . . . . 4
1.2.3 ResearchGap . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5
1.3 ResearchObjectives. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5
2 SignalandExploratoryDataAnalysis 6
2.1 SignalStatistics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
2.1.1 Log-returnasaHigh-PassFilter . . . . . . . . . . . . . . . . . . . . 6
2.1.2 SpectralandAutocorrelationAnalysis . . . . . . . . . . . . . . . . 7
2.1.3 ErgodicityChecking . . . . . . . . . . . . . . . . . . . . . . . . . . 8
2.1.4 Karhunen-LoèveandEmpiricalModeDecomposition . . . . . . . . 10
2.1.5 Time-FrequencyAnalysisviaContinuousWaveletTransform. . . . 13
2.2 ExploratoryDataAnalysis . . . . . . . . . . . . . . . . . . . . . . . . . . . 15
2.2.1 MarketversusStockBetas&CAPMrisk. . . . . . . . . . . . . . . 15
2.2.2 Market-level trenddrawdown&Regimeclassification . . . . . . . . 15
2.2.3 Vol-regimeZ-scores&Volatilitytermstructure . . . . . . . . . . . 16
2.2.4 MeanReversion&TrendConsistency. . . . . . . . . . . . . . . . . 18
2.2.5 Stock-LevelTrendOscillatorsandBoundedPercentileMetrics . . . 20
2.2.6 RSIvs.Pricevs.StochasticMomentum. . . . . . . . . . . . . . . . 21
2.2.7 StockMarketStructureandVolatilityCompression . . . . . . . . . 23
2.2.8 Stock-LevelTrendStrengthandReliability . . . . . . . . . . . . . . 25
2.2.9 RiskRegimeAnalysis. . . . . . . . . . . . . . . . . . . . . . . . . . 26
2.3 TimeSeriesDecompositionAnalysis . . . . . . . . . . . . . . . . . . . . . 29
2.3.1 Motivation. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29
2.3.2 Discussion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 33
3 StatisticalHypothesisTestingandInference 34
3.1 Methodology . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 34
SOICT-HUST
StockPriceAnalysis 2
3.1.1 Group1:TimeSeriesCharacteristics . . . . . . . . . . . . . . . . . 34
3.1.2 Group2:TechnicalAnalysisandMicrostructure . . . . . . . . . . . 35
3.1.3 Group3:MarketIndexRelationships . . . . . . . . . . . . . . . . . 35
3.2 HypothesisTestingResults . . . . . . . . . . . . . . . . . . . . . . . . . . . 35
3.2.1 Group1:TimeSeriesCharacteristicsofFPT. . . . . . . . . . . . . 36
3.2.2 Group2:TechnicalAnalysisandMicrostructureSignals. . . . . . . 37
3.2.3 Group3:MarketIndexRelationships . . . . . . . . . . . . . . . . . 38
3.3 ConclusionandModelingImplications . . . . . . . . . . . . . . . . . . . . 38
3.3.1 SummaryofStatisticalFindings . . . . . . . . . . . . . . . . . . . . 38
3.3.2 ImplicationsforFeatureSelection . . . . . . . . . . . . . . . . . . . 39
3.3.3 ImplicationsforForecastingModels . . . . . . . . . . . . . . . . . . 39
4 ForecastingModelDevelopment 40
4.1 ProblemFormulation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 40
4.2 DataRepresentationandFeatureEngineering . . . . . . . . . . . . . . . . 41
4.2.1 MarketFeatures. . . . . . . . . . . . . . . . . . . . . . . . . . . . . 41
4.2.2 StockFeatures . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42
4.2.3 DataProcessing&Pipeline . . . . . . . . . . . . . . . . . . . . . . 44
4.3 ModelSelection . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 48
4.3.1 SummaryofSelectedFeatures . . . . . . . . . . . . . . . . . . . . . 48
4.3.2 ClassicalModels . . . . . . . . . . . . . . . . . . . . . . . . . . . . 48
4.3.3 MachineLearningModels . . . . . . . . . . . . . . . . . . . . . . . 49
4.3.4 DeepLearningModels . . . . . . . . . . . . . . . . . . . . . . . . . 50
4.4 HyperparameterOptimization . . . . . . . . . . . . . . . . . . . . . . . . . 51
4.4.1 OptimizationMethodology. . . . . . . . . . . . . . . . . . . . . . . 51
4.4.2 HyperparameterSearchSpaces . . . . . . . . . . . . . . . . . . . . 51
4.4.3 SummaryofSearchSpaceParameters. . . . . . . . . . . . . . . . . 52
5 ModelEvaluationandConclusion 54
5.1 ForecastingPerformanceEvaluation. . . . . . . . . . . . . . . . . . . . . . 54
5.1.1 ExperimentalSetup. . . . . . . . . . . . . . . . . . . . . . . . . . . 54
5.1.2 EvaluationMetrics . . . . . . . . . . . . . . . . . . . . . . . . . . . 55
5.1.3 ForecastingResults . . . . . . . . . . . . . . . . . . . . . . . . . . . 55
5.2 StatisticalReliabilityofForecasts . . . . . . . . . . . . . . . . . . . . . . . 56
5.2.1 BootstrapConfidenceIntervals . . . . . . . . . . . . . . . . . . . . 56
5.2.2 Diebold–MarianoTest . . . . . . . . . . . . . . . . . . . . . . . . . 57
5.3 Conclusion. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 58
5.4 FutureWork. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 58
SOICT-HUST
Stock Price Analysis
3
Chapter 1
Introduction
1.1 Abstract
Stock price forecasting is challenging because financial time series are noisy, non
stationary, heavy-tailed, and affected by time-varying volatility. This study analyzes and
forecasts Vietnamese stock market data, focusing on FPT as the main stock and using
VNINDEX and VN30 as market references. The goal is not to guarantee profitable trad
ing decisions, but to build a statistically grounded framework for understanding market
behavior and comparing forecasting models.
The dataset is collected from VNStock and includes daily OHLCV data for FPT,
VNINDEX, and VN30. From the raw data, we construct log returns, lagged returns,
rolling statistics, volume-based features, technical indicators, and market-related vari
ables. We first conduct exploratory data analysis and statistical hypothesis tests, includ
ing stationarity, normality, autocorrelation, volatility clustering, Granger causality, vari
ance comparison, and market beta analysis. Then, we compare classical models such as
ARIMA/SARIMAX and GARCH-type models with machine learning models, including
XGBoost, MLP, LSTM, and GRU.
The empirical results show that FPT closing prices are non-stationary, while log re
turns are stationary. FPT returns also exhibit heavy tails and strong volatility clustering,
suggesting the need for robust metrics and volatility-aware features. In addition, FPT has
significant exposure to VNINDEX, showing that market-level movement is important for
explaining individual stock behavior. Model performance is evaluated using MAE, RMSE,
MAPE, directional accuracy, Diebold–Mariano tests, bootstrap confidence intervals, and
multi-horizon forecasting.
SOICT-HUST
Stock Price Analysis
4
1.2 Related Works
1.2.1 Classical Statistical Approaches
Classical time-series models such as ARIMA are widely used as interpretable base
lines for forecasting. The Box–Jenkins framework provides a systematic way to model
autoregressive and moving-average structures in stationary or differenced time series [1].
Since raw stock prices are often non-stationary, unit-root testing is necessary before model
fitting. The Dickey–Fuller test provides a standard method for detecting unit roots and
deciding whether differencing is required [2].
Financial returns also commonly exhibit volatility clustering, where large price changes
tend to be followed by large changes. Engle introduced the ARCH model to capture
conditional heteroskedasticity [3], and Bollerslev extended it into the GARCH model by
including past conditional variances [4]. These models are important because stock-market
risk is often time-varying rather than constant.
Market relationship analysis is another important part of financial modeling. The Effi
cient Market Hypothesis suggests that historical prices alone may have limited predictive
power in efficient markets [5]. Meanwhile, the Capital Asset Pricing Model measures sys
tematic market exposure through beta [6]. In this project, these ideas motivate the use of
VNINDEX and VN30 as market-level references when analyzing FPT.
1.2.2 Machine Learning in Finance
Machine learning models are often used in financial forecasting because they can cap
ture nonlinear relationships among features. XGBoost is a scalable gradient boosting
method that performs well on structured tabular data and can model complex feature
interactions [7]. Therefore, it is suitable for combining OHLCV variables, technical indi
cators, volume features, and market-index returns.
Deep learning models are also useful for sequential data. LSTM networks were designed
to learn long-term dependencies and reduce the vanishing-gradient problem in recurrent
neural networks [8]. GRU is a simpler gated recurrent architecture that can also capture
temporal dependencies with fewer parameters [9]. In this study, LSTM and GRU are used
to test whether recurrent models can improve short-horizon stock forecasting.
However, complex models can easily overfit noisy financial data. Therefore, model
comparison should not rely only on one error metric. The Diebold–Mariano test provides
a formal way to test whether two forecasting models have significantly different predictive
accuracy [10]. This project therefore evaluates models using both standard metrics and
statistical comparison tests.
SOICT-HUST
Stock Price Analysis
5
1.2.3 Research Gap
Many studies focus either on statistical testing or on predictive modeling. Pure statis
tical approaches are interpretable but may miss nonlinear patterns, while pure machine
learning approaches may ignore important properties such as non-stationarity, heavy tails,
and heteroskedasticity. This project bridges that gap by combining hypothesis testing, sig
nal processing, and predictive modeling in one workflow. Statistical tests guide feature
construction and model selection, while machine learning models are used to evaluate
predictive performance.
1.3 Research Objectives
The core objectives of this study are threefold:
• Statistical Testing: To identify key statistical properties of FPT, VNINDEX, and
VN30, including non-stationarity, heavy-tailed returns, autocorrelation, volatility
clustering, volume-related risk, and market dependence.
• Feature Engineering: To transform raw OHLCV data into useful predictive fea
tures, including log returns, lagged returns, rolling statistics, technical indicators,
market variables, and decomposition-based components.
• ModelConstruction and Evaluation:TobuildandcompareARIMA/SARIMAX,
GARCH-typemodels,XGBoost,MLP,LSTM,andGRUusingMAE,RMSE,MAPE,
directional accuracy, Diebold–Mariano tests, bootstrap confidence intervals, and
multi-horizon forecasting.
This study aims to answer the following questions:
1. Are FPT prices and log returns stationary?
2. Do FPT returns exhibit heavy tails or volatility clustering?
3. Do volume, technical indicators, VNINDEX, and VN30 provide useful predictive
information?
4. Do machine learning and deep learning models outperform classical statistical base
lines?
5. Which feature groups contribute most to forecasting performance?
SOICT-HUST
Stock Price Analysis
6
Chapter 2
Signal and Exploratory Data Analysis
2.1 Signal Statistics
2.1.1 Log-return as a High-Pass Filter
In financial time-series analysis, the continuous compounding return, or log-return, is
defined as:
rt = log Pt
Pt−1
= log(Pt)−log(Pt−1)
(2.1)
where Pt is the asset price at time t. By defining the log-price as pt = log(Pt), the log
return can be expressed simply as the first difference of the log-price series: rt = pt−pt−1.
From a signal processing perspective, this differencing operation acts as a linear time
invariant (LTI) filter applied to the log-price signal. Let the impulse response of this filter
be h[n]. The output rt is the convolution of pt with h[n], where h[0] = 1, h[1] = −1, and
h[n] = 0 otherwise. Taking the discrete-time Fourier transform (DTFT) of the impulse
response yields the frequency response function H(eiω):
∞
H(eiω) =
n=−∞
h[n]e−iωn = 1 −e−iω
(2.2)
where ω ∈ [0,π] represents the angular frequency. Therefore, we calculate its squared
magnitude:
|H(eiω)|2 = (1 − e−iω)(1 − eiω)
= 2(1 −cosω)
(2.3)
Analyzing this magnitude response reveals the filter’s characteristics across different
frequencies:
• Low Frequencies (ω → 0): As the frequency approaches zero, cos(ω) → 1, which
means |H(eiω)|2 → 0. Low frequencies represent slow-moving, long-term macroe
conomic trends in the stock price. The filter heavily attenuates these components,
SOICT-HUST
Stock Price Analysis
7
effectively removing the non-stationary “drift” or random walk component of the
price series.
• High Frequencies (ω → π): As the frequency approaches the Nyquist limit (π),
cos(ω) → −1, which means |H(eiω)|2 → 4. High frequencies represent rapid, day
to-day fluctuations and market microstructure noise. The filter not only retains but
amplifies these high-frequency components.
Therefore, transforming raw prices into log-returns systematically strips away the long
term trend (low frequencies) while preserving the short-term volatility (high frequencies).
Figure 2.1: Raw stock visualization
Figure 2.2: Log-return transformation
2.1.2 Spectral and Autocorrelation Analysis
To thoroughly understand the cyclical behavior and temporal dependencies of the
financial time series, we utilize both frequency-domain and time-domain analytical tech
niques.
SOICT-HUST
Stock Price Analysis
8
Power Spectral Density (PSD)
In the frequency domain, the Power Spectral Density describes how the variance (or
power) of a time series is distributed across different frequency components. To estimate
the PSD of the log-returns robustly, we apply Welch’s method [11]. Unlike a standard
periodogram, which can be highly noisy, Welch’s method divides the time series into over
lapping segments, computes a modified periodogram for each segment, and then averages
these estimates. This averaging significantly reduces the variance of the power spectrum
estimate.
Given a sampling frequency of fs = 1 sample per trading day, the frequency f repre
sents cycles per day. The reciprocal of the frequency yields the period T in days (T = 1/f).
In financial markets, distinct periodicities often emerge due to trading schedules and
macroeconomic reporting cycles. By plotting the PSD on a logarithmic scale, we can
identify these latent seasonalities. For instance, a peak at f ≈ 0.20 corresponds to a 5-day
cycle (weekly trading patterns), while a peak at f ≈ 0.047 corresponds to a 21-day cycle
(typical number of trading days in a month). We utilize a segment length of 252 days to
capture a full trading year of low-frequency dynamics.
Autocorrelation and Partial Autocorrelation
While the PSD identifies cyclical frequencies, the Autocorrelation Function (ACF) and
Partial Autocorrelation Function (PACF) are time-domain tools essential for identifying
the structural memory of the series.
The ACF measures the linear dependence between the current value of the series, rt,
and its past value at lag k, rt−k. It is defined as:
ρk = Cov(rt,rt−k)
Var(rt)
(2.4)
The ACF plot displays ρk against the lag k. The PACF, on the other hand, isolates the
direct correlation between rt and rt−k by removing the linear influence of the intermediate
lags (rt−1,rt−2,…,rt−k+1). Mathematically, it is the coefficient ϕkk in the autoregression
of rt on its k past values. If the PACF cuts off abruptly after lag p while the ACF decays
gradually, it indicates an AR(p) signature.
2.1.3 Ergodicity Checking
A fundamental assumption in time-series forecasting is ergodicity, which implies that
the time average of a single realization converges to the ensemble average (the expected
value across all possible realizations) as the sample size grows. To empirically verify ergod
icity, we must approximate an ensemble from a single observed realization. In this study,
SOICT-HUST
Stock Price Analysis
9
Figure 2.3: Auto and Partial-Auto Corre
lation plot for log-returns of VNINDEX
Figure 2.4: Power Spectral Density in fre
quency domain
we utilize two distinct simulation approaches: the K-Blocks method and the Stationary
Bootstrap method.
Method 1: The K-Blocks Approach
The K-blocks method partitions the single observed time series into contiguous seg
ments to simulate parallel universes. Let the time series of log-returns be denoted as
{r1, r2, …,rN}. We truncate the series slightly to length N′ such that it is perfectly
divisible by a chosen integer K, resulting in a block size of M = N′/K.
We construct an ensemble matrix where each row k represents a pseudo-independent
realization (block), and each column t represents an intra-block time step:
r1
R=
rM+1
.
.
.
r2
rM+2
.
.
.
. . .
rM
. . . r2M
. . .
.
.
.
r(K−1)M+1 r(K−1)M+2 … rKM
At each step t ∈ [1,M], the ensemble mean µ(E)
t
and ensemble variance σ2(E)
(2.5)
are calculated
down the columns across the K blocks. If the process is ergodic, these ensemble metrics
should stabilize and closely track the expanding time-average metrics of the original series.
t
Method 2: Stationary Bootstrap Approach
Financial returns often exhibit local dependencies (such as volatility clustering) that
are destroyed by simple random sampling. To address this, we apply the Stationary Boot
strap [12]. Unlike a standard block bootstrap that uses fixed-length blocks, the stationary
bootstrap resamples blocks of random lengths drawn from a geometric distribution with
expected length λ. This randomization ensures that the resulting synthetic paths remain
strictly stationary.
We generate B synthetic time-series paths, each of length N, to form a robust ensemble.
SOICT-HUST
Stock Price Analysis
10
Let ˜rb,t represent the return at time t for the b-th bootstrapped path. The bootstrapped
ensemble mean and variance at each time step t are given by:
B
µ(B)
t
= 1
B
B
b=1
˜rb,t, σ2(B)
t
= 1
B−1
b=1
2
˜
rb,t − µ(B)
t
(2.6)
Because this method preserves the full temporal length N rather than truncating into
sub-blocks, it provides a macro-level view of convergence.
Evaluation of Ergodicity
We visually compare the ensemble averages against the expanding time averages. The
global mean µ and global variance σ2 serve as the asymptotic baselines. It can be observed
(figure 2.5) that while the process is approximately ergodic w.r.t the mean (first moment),
the variance depends on time step (heteroskedasticity), constrasting the homoskedasticity
assumption of classic regression models.
Figure 2.5: Time-averaged mean & variance versus. ensemble statistics
2.1.4 Karhunen-Loève and Empirical Mode Decomposition
To isolate meaningful structural patterns from random market fluctuations, we con
duct advanced signal decomposition techniques including Karhunen-Loève (K-L) Expan
sion [13] to analyze the variance structure across fixed time windows and Empirical Mode
SOICT-HUST
Stock Price Analysis
11
Decomposition (EMD), adaptively separating frequency components across the entire se
ries.
Karhunen-Loève Expansion and Volatility Clustering
The Karhunen-Loève Expansion is a continuous analogue to Principal Component
Analysis (PCA). It projects a stochastic process onto a set of orthogonal basis func
tions, or eigenfunctions, ordered by the amount of variance they explain. When applied
to rolling 20-day windows (representing a standard trading month), the K-L expansion
reveals distinct differences in the predictability of directional returns versus market risk.
As shown in the cumulative explained variance analysis, raw returns exhibit behavior
consistent with the Efficient Market Hypothesis (EMH). The variance explained by the
principal components of raw returns grows linearly, indicating a highly noisy process dom
inated by idiosyncratic shocks where no small set of components can adequately capture
the total variance. Meanwhile, the squared returns—acting as a proxy for volatility—
demonstrate pronounced volatility clustering. The first few components capture a signif
icantly larger share of the variance (crossing the 50% threshold by the 6th component),
confirming that market risk contains a strong, low-dimensional underlying structure.
The first few eigenfunctions exhibit shapes that resemble well-known signal-processing
filters:
• PC1: A smooth, predominantly positive loading across the entire window, acting
as a low-frequency mode that captures the overall level of returns or volatility.
• PC2: A single sign change within the window, contrasting observations from the
first and second halves of the month and representing an intermediate-frequency
mode.
• PC3: Multiple sign changes and oscillatory behavior, corresponding to higher
frequency fluctuations within the window.
The eigenfunctions for volatility are notably smoother than those of raw returns, further
emphasizing that market volatility concentrates in clusters and evolves more predictably
than market direction.
Empirical Mode Decomposition (EMD)
While the K-L expansion is constrained by a fixed window size, Empirical Mode De
composition (EMD) is a fully data-driven approach that assumes no a priori basis func
tions [14]. EMD utilizes a sifting process to decompose the non-linear and non-stationary
log-return series into a finite set of Intrinsic Mode Functions (IMFs). Each IMF represents
a simple oscillatory mode with a distinct, locally defined frequency.
SOICT-HUST
Stock Price Analysis
12
In our experiment, the EMD applied to log-returns yields 16 distinct IMFs, effectively
separating the market’s temporal dynamics:
• High-Frequency Modes (IMF 1–4): These components contain the highest
frequencies and lowest amplitudes. They primarily capture market microstructure
noise, daily bid-ask bounce, and rapid, transient shocks.
• Intermediate Cycles (IMF 5–12): These modes represent medium-term market
memory, capturing weekly, monthly, and quarterly trading cycles, as well as the
momentum effects driven by institutional rebalancing and earnings reports.
• Low-Frequency Trends (IMF 13–16): The final IMFs and the residual represent
the slow-moving macroeconomic drift. They reflect long-term structural changes in
the Vietnamese economy and broad market cycles spanning multiple years.
Figure 2.6: K-L expansion of Raw log-return signal
Figure 2.7: Some IMF modes decomposed
SOICT-HUST
Stock Price Analysis
13
2.1.5 Time-Frequency Analysis via Continuous Wavelet Trans
form
To capture localized time-frequency dynamics, we apply the Continuous Wavelet
Transform (CWT). Unlike the Fourier transform, which uses infinite sine waves, CWT
uses localized “wavelets” that expand and contract to measure both when and at what
frequency structural shifts occur [15].
Using the Morlet wavelet as the mother wavelet, the CWT of the log-return series rt
at scale a and translation (time) b is defined as the convolution:
W(a,b) = 1
√
a
∞
−∞
rtψ∗ t−b
a dt
where ψ∗ is the complex conjugate of the Morlet wavelet function. The scale a is inversely
proportional to frequency: low scales correspond to high-frequency, short-term market
noise, while high scales correspond to low-frequency, long-term macroeconomic trends.
The localized variance, or wavelet power, is calculated as the squared magnitude of the
coefficients: P(a,b) = |W(a,b)|2 [16].
Interpretation of the Scalogram
The CWT Scalogram visualizes the wavelet power across time and scales, with some
insights regarding the return dynamics:
• Volatility Cascades: The bright regions (pink/light orange) representing high
wavelet power do not occur as horizontal bands (which would imply a constant cy
cle). Instead, they manifest as vertical “cones” or pillars. This indicates that major
market shocks are localized in time and cascade across multiple scales simultane
ously, disrupting both short-term trading patterns and long-term momentum.
• Short-Term Dominance during Crises: The highest power concentrations con
sistently appear at the top of the y-axis (scales a < 50). This confirms that during
periods of extreme market turbulence, a larger fraction of the local variance is con
centrated in short-term fluctuations.
Dominant Time Scale Dynamics
To further quantify these shifts, we extract the dominant time scale at each time step
t, defined as the scale that maximizes the wavelet power:
a∗
t = arg max
a
P(a,t)
The dominant scale frequently oscillates between two extremes:
SOICT-HUST
Stock Price Analysis
14
• High-Scale Regimes (a ≈ 150−250): During relatively calm market periods (e.g:
2018-2019), variance is primarily associated with slower-moving market dynamics
and persistent trends.
• Low-Scale Regimes (a < 50): During periods of uncertainty or sudden news
shocks (e.g: 2024-2025), the dominant scale abruptly collapses to the lower regis
ter. In these regimes, and price action is almost entirely dictated by short-term
fluctuations (for e.g: noise and reactionary trading).
Figure 2.8: Wavelet transformation
Figure 2.9: Dominant time scales plot
SOICT-HUST
Stock Price Analysis
15
2.2 Exploratory Data Analysis
2.2.1 Market versus Stock Betas & CAPM risk
Detailed formulation: 4.12 and 4.7
Analyzing the rolling Beta (figure 2.10) insights into FPT’s risk profile over time:
• Defensive Baseline: The empirical distribution of the 20-day rolling Beta reveals
a median value of 0.85. Because this is less than the market baseline (β = 1.0), FPT
is generally characterized as a defensive asset. On average, it exhibits less volatility
than the broader VNINDEX, cushioning portfolios during market downturns.
• Regime-Dependent Volatility: Despite its defensive median, the time-series plot
of the rolling Beta demonstrates significant non-stationarity. FPT’s systematic risk
fluctuates significantly, dropping near 0.0 (highly idiosyncratic movement discon
nected from the market) and occasionally spiking above 2.5 (extreme systematic
sensitivity). This indicates that FPT’s market exposure is highly regime-dependent;
during specific macroeconomic shocks or sector-wide momentum shifts, FPT can
temporarily abandon its defensive nature and behave as a highly aggressive, high
beta stock.
2.2.2 Market-level trend drawdown & Regime classification
Detailed formulation: 4.7 and 4.5
The VNINDEX serves as the primary proxy for Vietnamese market sentiment. By
evaluating the index’s historical drawdowns alongside its momentum-based trend strength,
we can explicitly categorize the market into distinct behavioral regimes, providing vital
context for downstream predictive algorithms.
Analyzing the historical trajectory of the VNINDEX through these lenses (Figure
2.11) reveals several structural insights:
• Cyclical and Severe Drawdowns: The drawdown dynamics illustrate a market
susceptible to deep, systemic corrections. The visualization clearly isolates two major
crisis periods: the acute, rapid crash in early 2020 (approaching a-40% drawdown)
and the prolonged, grinding bear market of 2022, which reached similar depths of
capital destruction.
• Momentum and Regime Shifts: The Trend Strength indicator effectively trans
lates moving-average crossovers into a continuous oscillator, cleanly segmenting
the timeline. Prolonged positive zones (such as late 2020 through 2021) represent
sustained bull regimes characterized by strong, persistent upward momentum and
SOICT-HUST
Stock Price Analysis
16
Figure 2.10: Illustration of FPT CAPM risk & Beta distribution
highly rapid recoveries from minor dips. Conversely, the deep negative troughs per
fectly capture the structural bear regimes.
• Transitional Market Phases: Beyond just bull and bear markets, the regime
classification identifies periods of macro-level indecision. During years like 2019 and
stretches of 2024, the trend strength oscillates tightly around the zero line, repre
senting sideways, low-momentum environments where mean-reversion dynamics are
more likely to dominate over trend-following momentum.
2.2.3 Vol-regime Z-scores & Volatility term structure
The visualizations are in Figure 2.12.
Volatility Term Structure Dynamics
Formulation: see 4.8, for fat tail figure, vol_regime_z is just Z-score of vol_20.
The 5-day volatility acts as a highly sensitive, noisy proxy for immediate market
reaction. As seen in the time-series plot, there are multiple periods where the 5-day
volatility violently spikes and breaks away from the smoother 20-day baseline. These
SOICT-HUST
Stock Price Analysis
17
Figure 2.11: Macro drawdown & Regime classification
massive divergences serve as early warning signals of acute market panic and sudden
liquidity shocks.
Fat-Tail Risk and Z-Score Distribution
The empirical distribution of Vol Z-scores provides a crucial warning regarding market
risk assumptions:
• Pronounced Positive Skew (Asymmetry): The histogram is not normally dis
tributed; the majority of the data clumps to the left of the mean. This confirms
that the stock market spends the vast majority of its time in a calm, low-volatility
state, often associated with slow, grinding upward trends.
• TheFat Right Tail: The most critical insight from the distribution is the existence
of a “fat” right tail. A significant density of observations extends far beyond the +2
standard deviation “panic threshold,” with extreme events reaching +4 or even +6
standard deviations, proving that extreme market shocks happen far more frequently
than standard Gaussian statistics would predict.
SOICT-HUST
Stock Price Analysis
18
Figure 2.12: Vol term structure & Fat tail analysis
2.2.4 Mean Reversion & Trend Consistency
The visual analysis relies on three primary engineered indicators:
• Momentum Z-Score: The 20-day price return is standardized using a 20-day
rolling mean and rolling standard deviation. This normalizes the momentum into
a Z-score, where values exceeding ±2 standard deviations theoretically indicate
statistical anomalies (overbought or oversold conditions).
• Local Drawdown: Measures the percentage drop from the highest close observed
over the previous 20 days, capturing the severity of localized market breakdowns.
Insights and Market Dynamics
Based on the generated plots, several insights regarding the market structure can be
observed:
Efficacy of Mean Reversion Signals: Figure 2.13 demonstrates that the benchmark
index exhibits strong mean-reverting tendencies on a 20-day horizon. The momentum
SOICT-HUST
Stock Price Analysis
19
Z-score frequently oscillates between the +2 (overbought) and −2 (oversold) thresholds.
Notably, extreme breaches of the −2Z lower bound (e.g., the liquidity shock in early
2020 and the sustained downtrend in late 2022) identify short-term exhaustion points.
These extreme readings consistently precede sharp corrective rallies or structural market
bottoms, validating the Z-score as a robust predictive feature for contrarian modeling
strategies.
Trend Instability and Drawdown Correlations: Figure 2.14 illustrates the relation
ship between Trend Consistency and Local Breakdowns. The 20-day win rate is highly
volatile, frequently crossing the 0.5 (50%) baseline, indicating that short-term price action
is highly stochastic and rarely sustains prolonged unidirectional streaks. Furthermore, pe
riods of severe 20-day local drawdowns (indicated by the deep red shaded regions) strongly
correlate with trend consistency plunging well below the 0.4 level. This suggests that ma
jor market corrections in this index are driven by persistent negative daily returns (a
sustained lack of “up days”) rather than isolated, single-day crash events.
Figure 2.13: Mean reversion signals: Momentum Z-scores
Figure 2.14: Trend consistency with Local breakdowns
SOICT-HUST
Stock Price Analysis
20
2.2.5 Stock-Level Trend Oscillators and Bounded Percentile Met
rics
To extract multi-scale momentum signals and identify trend reversals for the asset, we
analyze the structural behavior of the Moving Average Convergence Divergence (MACD)
indicator and its normalized variants across a multi-year horizon (2021 to mid-2025). The
feature formulations are in 4.17.
Momentum Histogram Dynamics and Velocity Acceleration
The MACD histogram (Ht = MACDt−Signalt) isolates the distance between the core
trend line and its slower moving average. Rather than mapping the trend direction itself,
the histogram acts as a proxy for the second derivative of price, capturing market velocity
and directional acceleration.
An analysis of the green (bullish) and red (bearish) momentum waves across the
timeline yields several insights for short-horizon forecasting:
• Leading Indicator Properties (Momentum Deceleration): The histogram
(figure 2.15) serves as a leading indicator for structural price turning points. As
seen in the mid-2023 and mid-2024 expansions, the green bullish histogram blocks
reach their peak amplitude and begin to slope downward before the absolute price
peak is reached and before a formal bearish crossover trigger occurs. This shrinking
of the green bars signifies a loss of buying velocity (negative acceleration).
• High-Frequency Mean Reversion via Histogram Percentiles: In Figure 2.16,
the teal line representing the 20-day histogram percentile (hist_percentile_20)
exhibits a much higher cycling frequency than the smoother MACD line percentile.
Because the histogram shifts overextended values back toward zero rapidly, its per
centile tracks localized overbought and oversold states with high sensitivity.
Figure 2.15: MACD vs. Signal line
SOICT-HUST
Stock Price Analysis
21
Figure 2.16: Bounded Percentiles Analysis
2.2.6 RSI vs. Price vs. Stochastic Momentum
To evaluate structural momentum exhaustion and directional reliability, we analyze
multi-scale oscillator dynamics through Relative Strength Index (RSI) and Stochastic
Momentum tracking, alongside engineered indicators for signal quality and divergence
(See 4.2.2).
Time-Frequency Duality in Oscillators (RSI vs. Stochastic %K)
The Figure 2.17 demonstrates a clear time-frequency structural decoupling between
the two indicators:
• High-Frequency Volatility of Stochastic %K: The Stoch_K indicator (orange
line) acts as a high-frequency oscillator, rapidly oscillating across its entire domain
[0, 100]. It reaches extreme overbought (≥ 80) and oversold (≤ 20) thresholds fre
quently. This behavior captures localized closing price placement relative to the
recent 14-day high-low range, making it highly sensitive to short-horizon mean
reversion waves but noisy during sustained trends.
• Low-Pass Smoothing of RSI(14): In contrast, the RSI_14 indicator (blue line)
functions as a smoother, lower-frequency momentum tracker. It filters out minor
microstructure noise and remains bounded within intermediate ranges for extended
periods.
• Conjoint Extremes as Regime Exhaustion Triggers: A powerful predictive
feature occurs when both oscillators reach extreme saturation thresholds simul
taneously. For instance, in mid-2023, both the blue and orange lines hit extreme
upper boundaries (> 80), flagging a robust, highly overextended bullish regime.
Conversely, when both collapse simultaneously to the lower register (as seen in deep
market troughs), it indicates high-probability capitulation zones where selling ve
SOICT-HUST
StockPriceAnalysis 22
locity is statisticallyexhausted.Otherwise,whentheirdirections contradict, the
momentumhasbeenweakenedandlessreliable.
MomentumReliabilityandStructuralDivergenceMetrics
Thefeaturemomentum_qualityiscomputedas:LetCROC,t∈{−1,0,1}representthe
ROCdirectionalconsistencyindicatorattimet,derivedfromthealignmentofshort-and
medium-termhorizons.LetMc,tdenotethemomentumcompositescore,andσMc,t signify
its20-dayrollingstandarddeviation.Toaccountforinstitutional liquidityvalidation,the
volumeratioVratio,t isdefinedas:
Vratio,t= Vt
1
20
19
i=0
Vt−i
(2.7)
whereVt istherawtradingvolumeattimet.
Weutilizean indicator functionI(·) that evaluates to1 if the interior condition is
satisfiedand0otherwise.Toensureboundedstability fordownstreammachine learn
ing inputs, thecomposite rawscore is clippedwithinthedomain [0,1].Dependingon
theavailabilityofmarket liquiditydata,MQt is computedviaaconditional piecewise
framework:
MQt=
clip 0.3CROC,t+0.4I(|Mc,t|>σMc,t)+0.3I(Vratio,t>1) , ifvolumeisavailable,
clip 0.5CROC,t+0.5I(|Mc,t|>σMc,t) , ifvolumeisunavailable.
(2.8)
whereclip(x)=max(0,min(1,x)).
TheFigure2.18illustratesthestabilityandstructuralalignmentofthesemomentum
wavesusingtwoengineeredfeatures:momentum_qualityandRSI_price_divergence.
•RegimeValidationviaMomentumQuality:Themomentum_quality(navy
blue line) indicator rapidlymovesbetweendiscretestructural levels (0.0,0.5,1.0).
Whenitpinsat1.0, itconfirmsahigh-conviction,broad-basedtrend,givingdown
streammachine learningmodelsaclearsignal totrusttrend-followingrules.Con
versely,whenitplummetsto0.0, itwarnsthemodel thatthecurrentpriceaction
lacksvolumeorconsistency,signalinganoisyorrangingenvironment.
•DivergenceCaptureasaTurningPointPredictor:TheRSI_price_divergence
indicator(darkorangeline)quantifiesdirectionalmismatchesbetween14-dayprice
returnsand14-dayRSIadjustments.Itoperatesasadistinctphaseshifterstepping
between+1.0,0.0,and−1.0:
SOICT-HUST
Stock Price Analysis
23– Value of 0.0: Indicates complete structural agreement between price momen
tum and oscillator momentum.– Spikes to ±1.0: Represent hidden structural anomalies. As observed in late
2023 and early 2025, persistent jumps to extreme divergence thresholds appear
right before major trend corrections or trend reversals. Bullish divergence (< 0)
means decreasing sell force, whereas bearish divergence (> 0) indicates building
selling pressure.
Figure 2.17
Figure 2.18: RSI price divergence and momentum quality indicators
2.2.7 Stock Market Structure and Volatility Compression
To model localized phase transitions and anticipate explosive directional expansions,
we map the geometric boundaries of the asset’s price action. By establishing non-parametric
support and resistance channels based on rolling extrema over a lookback window of
w = 20 days, we extract structural covariates that capture breakout intensity and non
linear volatility cycles.
SOICT-HUST
Stock Price Analysis
24
Dynamic Market Structure and Breakout Intensity
The architectural envelope of the asset’s price is defined by the previous 20-day high
and low, constructing a dynamic horizontal corridor (figure 2.19). Analyzing the distri
bution of prices relative to these boundaries reveals distinct structural regimes:
• Structural Consolidation and Range Trading: From mid-2021 through mid
2023, the asset is trapped in a protracted horizontal baseline between $40 and $60.
During this phase, the price repeatedly tests the boundaries without sustaining a
breakout, keeping the range position feature oscillating symmetrically.
• Persistent Trend Extension: Beginning in mid-2023 and continuing throughout
2024, a powerful structural shift occurs. The price consistently overrides the trailing
resistance line, generating a series of consecutive upside breakouts
(breakout_high_strength_20 > 0, highlighted in lime green).
• Structural Breakdown Regime: In early 2025, after peaking near $130, the mar
ket structure flips abruptly. The price plunges through the trailing 20-day support
floor, triggering a severe downside breakdown phase
(breakout_low_strength_20 > 0, highlighted in crimson). This transition signals
some oversell and a fundamental breakdown of the medium-term uptrend.
The Volatility Squeeze and Equilibrium Expansion
A core principle of market microstructure is that periods of extreme volatility compres
sion act as precursors to structural expansions. To capture this phenomenon mathemati
cally, we track the normalized width of the local channel via the structure_tightness_20
metric, alongside its localized 60-day statistical deviation (tightness_z).
The time-series visualization (figure 2.20) of this tightness profile reveals predictive
behaviors:
• Coiled Energy States (The Squeeze): When the channel width collapses signif
icantly below its rolling average, the Z-score crosses the threshold of tightness_z <
−1.5. This triggers an extreme compression state, visualized as orange vertical
“Squeeze” blocks. These episodes represent temporary market equilibriums where
buyers and sellers are tightly matched.
• Predictive Anchor for Regime Shifts: Examining the alignment between the
two figures demonstrates that these extreme compression zones systematically pre
cede major, high-velocity price moves. Notable squeeze triggers appear in mid-2022,
late 2022, and late-2024 right before the massive structural breakout.
SOICT-HUST
Stock Price Analysis
25
Figure 2.19: Stock price structure with Resistance & Support (Zoomed in)
Figure 2.20: Stock price tightness
2.2.8 Stock-Level Trend Strength and Reliability
We evaluate FPT’s directional behavior by analyzing the first derivative of its expo
nential moving average (Trend Velocity) alongside its rolling non-parametric directional
persistence (Trend Consistency). The visualizations are in Figure 2.21.
Trend Velocity and Moving Average Derivatives
This feature models the momentum of the 20-day EMA using a 5-day (ema_slope_20_5)
and 10-day (ema_slope_20_10) lookback shift. Because these metrics calculate the rate
of change of an already smoothed baseline, they filter out high-frequency microstructural
noise, exposing the true velocity of structural shifts:
• Velocity Acceleration Regimes: The green shaded expansions show periods
where the fast velocity line rises above the zero threshold, indicating an accelerating
uptrend. Strong, sustained clusters of positive velocity are highly visible during the
structural bull markets of 2020-2022 and the aggressive extension phases throughout
2023-2024.
SOICT-HUST
Stock Price Analysis
26
• Asymmetric Downward Shocks:Thered shaded regions isolate periods of struc
tural deceleration and negative velocity. During major liquidity crises—such as the
early 2020 COVID crash, the late 2022 global bear market, and the sharp early
2025 reversal—the fast velocity line violently plunges below −0.010. This reveals
that downward velocity in FPT’s trend profile operates with much greater sudden
ness and intensity than upward acceleration.
• Crossover Signal Lags: The interaction between the fast blue line and the slow
orange dashed line acts as an effective internal momentum stabilizer. When the
fast velocity line crosses back toward the zero baseline while the slow line remains
extended, it provides downstream models with a clear indicator of momentum de
celeration prior to an actual price reversal.
Trend Consistency and Win-Rate Probability Bounds
This feature quantifies the structural persistence of the asset’s trajectory. By com
puting a rolling 20-day mean of positive daily return counts, this feature acts as a local
win-rate probability index, standardizing directional momentum into a strictly bounded
domain:
• Empirical Probability Boundaries: The win rate is highly non-random, oscil
lating reliably within a structured envelope between a lower floor of ∼ 20% and an
upper ceiling of ∼ 75%.
• Persistence as a Regime Anchor: During powerful structural trends, the win
rate stays pinned above the neutral 50% baseline for consecutive months. These
green-filled regimes reveal a market with strong positive drift, where daily gains
consistently outnumber daily losses, implying trend-following instead of mean re
versions.
• Capitulation Floors: Conversely, during macro market liquidations, the win rate
collapses down to the 20% to 30% threshold. These deep red troughs capture periods
of extreme oversold capitulation. Historically, when FPT’s 20-day win rate hits this
structural floor (as seen in mid-2018, early 2020, late 2021, and early 2025), it marks
a high-probability boundary where selling exhaustion is achieved.
2.2.9 Risk Regime Analysis
We evaluate FPT’s structural risk architecture by tracking the dynamics of the Average
True Range (ATR) alongside the Downside Volatility Ratio across the 2021–2025 timeline.
SOICT-HUST
Stock Price Analysis
27
Figure 2.21: Trend velocity & reliability
Intraday and Gap Risk Tracking via ATR Dynamics
The first plot 2.22 decouples the absolute price-normalized trading range (atr_pct)
from its short-term historical baseline (atr_rel). Analyzing the interaction between these
two metrics yields several insights:
• Scale-Normalized Risk Context: The absolute ATR percentage (black line)
demonstrates that FPT’s typical daily trading range shifts between a calm base
line of 1% to 2% of the stock price and extreme historical stress levels exceeding
5% to 6%. This slow-moving baseline shift captures structural changes in market
liquidity and capitalization over multiple years.
• Volatilty Expansion Shocks: The Relative ATR (orange line) normalizes these
shifts by evaluating the current 14-day window against a 20-day macro baseline.
Periods where the orange line violently pierces the +1.5 threshold (“High Volatility
Expansion”) capture sudden, high-velocity regime changes.
• Intraday Noise Separation: Notably, the relative ATR spikes can occur when the
absolute ATR percentage is low (e.g., mid-2023 and early 2024). This provides the
forecasting models with an early-warning breakout filter, flagging moments where
SOICT-HUST
Stock Price Analysis
28
trading ranges are expanding aggressively relative to the immediate past, indepen
dent of long-term price inflation.
Risk Asymmetry and Toxic Volatility Fractions
The second plot 2.23 calculates the proportion of total variance attributed exclusively
to negative return days (downside_vol_ratio).
• Pervasive Downside Dominance: The crimson line oscillates primarily within
an elevated channel between 0.4 and 0.7, revealing that FPT’s total volatility is
consistently heavily influenced by downward price action.
• The 0.7 Danger Zone Threshold: The horizontal dashed baseline at 0.7 delin
eates the “Danger Zone” where downside volatility compromises more than 70%
of the asset’s total variance. The red-shaded peaks (“Dominant Downside Risk”)
mark moments of severe structural decay. These episodes map to major market ca
pitulations—specifically in mid 2021, late-2022, and a dense, persistent cluster in
early-to-mid 2025.
Figure 2.22: Gaps/Intra-day risk by ATR%
Figure 2.23
SOICT-HUST
Stock Price Analysis
29
2.3 Time Series Decomposition Analysis
2.3.1 Motivation
Time series decomposition was investigated as a preliminary step to determine whether
the FPT stock price series could be represented as a combination of interpretable compo
nents such as trend, seasonality, and irregular fluctuations. If stable trend and seasonal
structures existed, a decomposition-based hybrid forecasting framework could be con
structed by modeling each component separately and subsequently combining the fore
casts.
For an additive decomposition model, the observed series can be expressed as
yt = Tt +St +Rt
(2.9)
where Tt denotes the trend component, St denotes the seasonal component, and Rt
represents the residual or irregular component.
To evaluate the existence of meaningful seasonal patterns in FPT stock prices, several
decomposition techniques were examined, including Classical Seasonal Decomposition,
Seasonal-Trend Decomposition using Loess (STL), and Singular Spectrum Analysis (SSA).
Classical Decomposition and STL Analysis
To investigate whether the FPT stock price series contains meaningful seasonal pat
terns, both Classical Seasonal Decomposition and Seasonal-Trend Decomposition using
Loess (STL) were applied. Since financial time series rarely exhibit obvious periodic be
havior, multiple candidate seasonal periods were examined, including short-term cycles
(5, 10, and 20 trading days) and longer-term cycles (30, 60, and 120 trading days).
For the classical decomposition approach, the time series was decomposed into trend,
seasonal, and residual components under different seasonal assumptions. Figure 2.24 sum
marizes the decomposition results obtained using the seasonal_decompose method.
SOICT-HUST
Stock Price Analysis
30
Figure 2.24: Classical seasonal decomposition of the FPT stock price series under repre
sentative seasonal assumptions.
Visual inspection indicates that the extracted seasonal components are relatively weak
compared with the long-term trend and irregular fluctuations. Moreover, the seasonal
patterns vary considerably across different candidate periods and fail to exhibit consistent
recurring structures. These observations suggest that any apparent seasonality is likely
driven by random market movements rather than genuine deterministic cycles.
To provide a more quantitative assessment, STL decomposition was subsequently ap
plied and the seasonal strength metric proposed by Hyndman and Athanasopoulos was
computed:
FS = max 0,1− Var(Rt)
(2.10)
Var(St + Rt)
where values close to 1 indicate strong seasonality and values close to 0 indicate weak
seasonal effects.
The seasonal strength was evaluated across all candidate periods. Figure 2.25 presents
the resulting seasonal strength values obtained from STL decomposition.
SOICT-HUST
Stock Price Analysis
31
Figure 2.25: Seasonal strength (FS) estimated from STL decomposition across multiple
candidate seasonal periods.
The results consistently indicate weak seasonality. Across all tested periods, the sea
sonal strength remains below 0.2, implying that the extracted seasonal components ex
plain only a small proportion of the total variance in the series. Furthermore, no dominant
period emerges from the analysis, indicating the absence of a stable recurring cycle.
The findings obtained from both Classical Decomposition and STL lead to the same
conclusion: the FPT stock price series does not exhibit sufficiently strong deterministic
seasonality to justify decomposition-based forecasting. Consequently, traditional decom
position approaches were not considered suitable as the primary decomposition stage for
the proposed forecasting framework.
Singular Spectrum Analysis (SSA)
Because both Classical Decomposition and STL rely on predefined seasonal assump
tions, an additional experiment was conducted using Singular Spectrum Analysis (SSA),
a non-parametric decomposition technique capable of extracting latent structures without
requiring an explicit seasonal period.
SSA consists of four major steps:
1. Embedding the original time series into a trajectory matrix;
2. Applying Singular Value Decomposition (SVD) to the trajectory matrix;
3. Grouping singular components according to their contribution;
4. Reconstructing the series from selected eigentriples.
Unlike traditional decomposition methods, SSA does not assume a fixed trend-seasonal
structure and can identify low-frequency and high-frequency components directly from the
data.
SOICT-HUST
Stock Price Analysis
32
The decomposition results obtained from SSA were noticeably more informative than
those produced by Classical Decomposition and STL. The first reconstructed components
captured the dominant long-term dynamics of the stock price, while higher-order com
ponents represented short-term oscillations and stochastic fluctuations. In particular, the
reconstructed trend component appeared smoother and more stable than those obtained
from conventional decomposition techniques.
The decomposition reveals that a substantial proportion of the total variance is con
centrated within the first few principal components, while the remaining components
primarily capture noise-like behavior and localized fluctuations. This separation provides
a clearer representation of the underlying structure of the FPT stock price series.
From an exploratory perspective, SSA offers a more effective decomposition of financial
time series than methods based on predefined seasonal assumptions. However, despite
producing visually meaningful trend and noise separation, SSA does not reveal a stable and
interpretable seasonal mechanism. Therefore, SSA was primarily used as an exploratory
analytical tool rather than as a preprocessing stage for the final forecasting framework.
Limitations of SSA for Hybrid Forecasting
Although SSA produced visually meaningful decompositions, several limitations pre
vented its adoption as the decomposition stage of the final hybrid forecasting framework.
First, the reconstructed SSA components do not possess direct economic interpre
tation. Unlike classical trend and seasonal components, individual SSA components are
mathematical constructs derived from singular vectors and therefore cannot be readily
associated with specific market mechanisms.
Second, SSA decomposition is sensitive to several hyperparameters, particularly the
embedding dimension and component grouping strategy. Different parameter selections
may yield substantially different reconstructed components, reducing the robustness and
reproducibility of the decomposition.
Third, SSA may introduce information leakage when applied improperly in forecast
ing settings. Since decomposition is often performed using the entire available sample,
reconstructed components may inadvertently incorporate future information that would
not be available in a real-time forecasting environment. This can lead to overly optimistic
performance estimates.
Finally, despite producing a cleaner separation between low-frequency and high-frequency
structures, SSA did not reveal a stable and interpretable seasonal mechanism. The domi
nant components primarily reflected long-term trends and stochastic market fluctuations
rather than recurring seasonal cycles that could be modeled independently.
For these reasons, SSA was considered an exploratory analytical tool rather than a
preprocessing stage for the final forecasting architecture.
SOICT-HUST
Stock Price Analysis
33
2.3.2 Discussion
The decomposition experiments provide important insights into the statistical charac
teristics of the FPT stock price series.
Both Classical Decomposition and STL indicate that seasonal effects are weak and
unstable, as evidenced by seasonal strength values consistently below 0.2. These findings
suggest that deterministic seasonality contributes little to the overall variation of stock
prices.
SSA successfully identifies latent trend structures and separates noise components
more effectively than STL. However, the extracted components lack strong interpretability
and do not support a reliable decomposition-based forecasting framework.
Overall, the decomposition analysis suggests that FPT stock prices are primarily
driven by trend dynamics, stochastic fluctuations, and volatility clustering rather than
stable seasonal behavior. This observation is consistent with the stationarity tests, return
distribution analysis, and volatility analysis presented in previous sections. Consequently,
subsequent forecasting models focus on return-based representations, lagged features,
volatility-aware indicators, and machine learning methods instead of decomposition-based
hybrid forecasting approaches.
SOICT-HUST
Stock Price Analysis
34
Chapter 3
Statistical Hypothesis Testing and In
ference
This chapter validates the statistical assumptions behind the forecasting framework.
The analysis focuses on FPT stock, VN30, and VNINDEX, and is organized into three
main parts: methodology, hypothesis testing results, and modeling conclusions. The tests
help determine whether the data should be modeled in price or return form, whether
volatility requires special treatment, and which market or technical variables are statisti
cally useful for prediction.
3.1 Methodology
The hypothesis testing procedure follows a consistent decision rule. For each hypothe
sis, a null hypothesis H0 and an alternative hypothesis H1 are defined. The statistical test
produces a test statistic and a p-value. At the significance level α = 0.05, H0 is rejected
if p < 0.05; otherwise, there is insufficient evidence to reject H0.
The dataset contains daily market observations from 2020 to 2024. The main variables
used in this chapter are FPT closing price, FPT logarithmic return Rt = ln(Pt/Pt−1),
VN30 return, VNINDEX return, trading volume, and derived technical indicators. The
methodology is divided into three hypothesis groups, each presented in a separate sub
section below.
3.1.1 Group 1: Time Series Characteristics
The first group evaluates whether FPT prices and returns satisfy basic assumptions for
time series modeling. The Augmented Dickey–Fuller (ADF) test is used to check station
arity. The Jarque–Bera test evaluates normality and fat-tailed behavior. The Ljung–Box
Q-test checks whether returns contain significant linear autocorrelation. Engle’s ARCH
LM test is used to detect volatility clustering and conditional heteroskedasticity.
SOICT-HUST
Stock Price Analysis
35
These tests are important because forecasting models should not be fitted directly
to non-stationary price series without transformation. If returns are stationary but non
normal and heteroskedastic, the model should use returns as the target variable and
include volatility-aware components.
3.1.2 Group 2: Technical Analysis and Microstructure
The second group evaluates whether trading volume and technical indicators provide
statistically significant information. Granger causality is used to test whether volume
predicts future FPT returns. Levene’s test is used to compare variance between high
volume and normal-volume periods. Regression coefficient tests evaluate lagged-return
momentum. Two-sample tests evaluate whether RSI oversold signals produce abnormal
future returns, while variance tests evaluate whether Bollinger Bands squeezes are followed
by volatility expansion.
This group supports feature selection. A technical indicator should be treated cau
tiously if it does not show significant standalone predictive power. However, even a non
directional signal may still be useful if it helps identify higher-risk or higher-volatility
periods.
3.1.3 Group 3: Market Index Relationships
The third group evaluates the relationship between FPT and broad market indices.
A CAPM-style regression estimates the systematic risk coefficient β between FPT and
VNINDEX. Granger causality tests whether VN30 leads FPT returns. A coefficient com
parison test checks whether FPT reacts asymmetrically to market upturns and downturns.
A paired t-test evaluates whether FPT outperforms VN30 on average. Finally, a Pear
son correlation test evaluates whether VNINDEX trading volume is associated with FPT
volatility.
These tests determine whether market-index variables should be included as exogenous
predictors in later forecasting models such as ARIMAX, XGBoost, or hybrid models.
3.2 Hypothesis Testing Results
This section reports the empirical results for 15 hypotheses. To keep the chapter struc
ture clear, each hypothesis group is presented as a separate subsection: Group 1 for time
series characteristics, Group 2 for technical and microstructure signals, and Group 3 for
market-index relationships.
SOICT-HUST
Stock Price Analysis
36
3.2.1 Group 1: Time Series Characteristics of FPT
This group tests whether FPT price and return series are appropriate for statistical
forecasting. The results are summarized in Table 3.1.
Table 3.1: Hypothesis Tests: Time Series Characteristics
ID Hypothesis/Test
Stat
p-value Reject H0 Conclusion
1 Stationarity of FPT
price (ADF)
2 Stationarity of FPT re
turn (ADF)
3 Normality of FPT re
turn (Jarque–Bera)
4 Autocorrelation of re
turns (Ljung–Box)
5 Volatility
(ARCH LM)
clustering
−0.0514
0.954
−32.3825 < 0.001
885.13
10.50
162.26
<0.001
0.398
<0.001
No
Yes
Yes
No
Yes
FPT price is non-stationary
and contains a unit root.
FPT return is stationary
and suitable as a modeling
target.
Returns are non-normal
and fat-tailed.
No significant linear auto
correlation is detected.
Strong ARCH effect and
volatility clustering exist.
The results indicate that raw FPT prices are non-stationary and therefore unsuitable
as a direct target for predictive modeling. In contrast, FPT log returns exhibit clear
stationarity, making them more appropriate for forecasting tasks.
The distributional analysis further reveals strong deviations from normality, with the
Jarque–Bera test strongly rejecting the Gaussian assumption (p-value ≈ 0) and excess
kurtosis of 6.25 compared to the Gaussian benchmark of 3.0. These findings confirm that
FPT returns exhibit pronounced heavy-tailed behavior, implying a higher likelihood of
extreme market movements.
In terms of temporal dependence, the Ljung–Box test indicates that log returns ex
hibit no statistically significant linear autocorrelation (p-value > 0.05), suggesting limited
effectiveness of purely mean-based linear models in capturing return dynamics. However,
the Engle’s ARCH LM test strongly rejects the null hypothesis of homoskedasticity (p
value < 10−29), providing overwhelming evidence of pronounced volatility clustering. This
implies that the conditional variance of returns is time-varying and exhibits persistent
clustering over time, even in the absence of linear dependence in the mean process. Con
sequently, these results strongly motivate the use of GARCH-family models for effectively
modeling and forecasting volatility dynamics in FPT returns.
Given these characteristics, we recommend modeling frameworks that explicitly cap
ture volatility dynamics, such as GARCH-type models, rather than purely autoregressive
mean models such as ARIMA.
Finally, due to the heavy-tailed nature of the return distribution, the Mean Absolute
Error (MAE) is adopted as the optimization objective instead of RMSE. MAE is more
robust to extreme values and outliers, which are frequent in the FPT return series, and
therefore provides a more stable and representative training signal for model selection and
SOICT-HUST
Stock Price Analysis
37
hyperparameter tuning.
3.2.2 Group 2: Technical Analysis and Microstructure Signals
This group evaluates whether volume, momentum, RSI, and Bollinger Bands provide
useful information for forecasting. The results are summarized in Table 3.2.
Table 3.2: Hypothesis Tests: Technical Analysis and Microstructure
ID Hypothesis/Test
Stat
p-value Reject H0 Conclusion
6 Volume-price Granger
causality
7 Volume spikes and
volatility
8 Lag-1 return momen
tum
9 RSI oversold T+3 re
turn
10 Bollinger squeeze and
volatility
F-test min
0.129
Levene: 137.16 < 0.001
βLag1 = 0.0205 0.360
RSI< 30: 0.94% 0.214
Levene: 6.04
0.993
No
Yes
No
No
No
Trading volume does not
significantly predict future
FPT returns.
High-volume days have
about 4.70× higher return
variance.
Lag-1 return does not sig
nificantly predict next-day
return.
RSI oversold signals are not
statistically superior.
Squeeze periods do not sig
nificantly expand five-day
volatility.
The results indicate that most standalone technical indicators do not provide statisti
cally reliable predictive power for return direction. Trading volume does not Granger-cause
returns, lagged momentum signals are not statistically significant, and traditional oscil
lators such as RSI and Bollinger Band-based signals fail to exhibit consistent predictive
effects at the 5% significance level.
A more nuanced picture emerges when analyzing specific market conditions. Although
the RSI oversold condition (RSI < 30) is associated with higher subsequent T+3 average
returns (0.94% compared to 0.31% under normal conditions), this difference is not statis
tically significant due to limited sample size and high variance, suggesting that RSI-based
mean reversion signals should not be used in isolation.
Similarly, Bollinger Band squeeze conditions (defined as extremely narrow bandwidth
below the 5th percentile) do not lead to a statistically significant increase in subsequent
return volatility. Instead, FPT prices tend to remain in prolonged low-volatility regimes
following such compression periods before a clear directional move emerges.
In contrast, the most robust and economically significant result is observed in volume
based anomalies. Days with abnormal trading volume (exceeding 2× the 20-day moving
average) exhibit approximately 4.7 times higher return variance compared to normal days,
with a near-zero p-value. This finding strongly confirms that volume spikes are primarily
a risk and volatility signal rather than a direct predictor of return direction.
SOICT-HUST
Stock Price Analysis
38
Overall, these results suggest that technical indicators in this setting are more infor
mative for volatility and regime detection than for directional return forecasting.
3.2.3 Group 3: Market Index Relationships
This group examines whether VNINDEX and VN30 provide useful external informa
tion for explaining or forecasting FPT returns. The results are summarized in Table 3.3.
Table 3.3: Hypothesis Tests: Market Index Relationships
ID Hypothesis/Test
11 Systematic risk CAPM
beta
12 VN30 Granger causal
ity on FPT
13 Asymmetric market
impact
14 Outperformance ver
sus VN30
15 VNINDEX liquidity
and FPT volatility
Stat
β =0.9562
F-test min
∆β =0.0028
Paired t-test
p-value Reject H0 Conclusion
<0.001
0.070
0.966
0.003
Pearson r = 0.0228 0.156
Yes
No
No
Yes
No
FPT has a strong sys
tematic relationship with
VNINDEX (R2 =46.6%).
VN30 does not Granger
cause FPT returns at the
5% level.
FPT beta is symmetric in
market upturns and down
turns.
FPT significantly out
performs VN30 by about
+0.0755% per day.
Market liquidity is not sig
nificantly correlated with
FPT volatility.
The market relationship tests show that VNINDEX is an important explanatory vari
able for FPT. The CAPM beta is highly significant and close to one, meaning that FPT
moves strongly with the market. However, VN30 does not significantly lead FPT returns
at the 5% level, and FPT’s beta does not differ between bullish and bearish regimes. The
paired test shows that FPT significantly outperforms VN30 on average, while VNINDEX
volume does not explain FPT volatility.
3.3 Conclusion and Modeling Implications
The hypothesis testing results provide several important conclusions for the forecasting
framework developed in later chapters.
3.3.1 Summary of Statistical Findings
Across the 15 hypotheses, six results are especially important. First, FPT price is non
stationary, so raw prices should not be used directly as the main target without transfor
mation. Second, FPT returns are stationary, making them more appropriate for time series
modeling. Third, return normality is rejected, meaning that the model should account for
SOICT-HUST
Stock Price Analysis
39
fat tails and outliers. Fourth, volatility clustering is strongly significant, supporting the
inclusion of GARCH-type models or volatility-related features. Fifth, abnormal volume is
associated with substantially higher variance, making it useful for risk detection. Sixth,
VNINDEX has a strong and significant relationship with FPT returns, so market-index
variables should be considered as exogenous predictors.
3.3.2 Implications for Feature Selection
The tests suggest that not all technical indicators should be treated equally. RSI
oversold signals, Bollinger squeeze signals, lag-1 return momentum, and trading volume
alone do not show strong standalone predictive power for return direction. Therefore,
these variables should not be interpreted as reliable independent trading signals. They
may still be useful when combined with other variables in nonlinear models, but their
individual statistical evidence is weak.
In contrast, volume spikes and VNINDEX returns provide stronger evidence. Volume
spikes are useful for identifying high-volatility periods, while VNINDEX returns capture
systematic market movement. These variables should be prioritized in the feature engi
neering stage.
3.3.3 Implications for Forecasting Models
The modeling pipeline should use return-based targets rather than raw prices. Because
the return distribution is fat-tailed, robust objectives such as MAE, Huber loss, or quantile
loss are more appropriate than purely Gaussian-error assumptions. Since volatility clus
tering is significant, volatility-aware models such as GARCH, hybrid ARIMA–GARCH,
or machine learning models with rolling volatility features are justified.
For exogenous variables, VNINDEX should be included because it explains a substan
tial share of FPT return variation. VN30 may be included as a supplementary market
feature, but the evidence for a direct leading effect is weaker. Overall, the statistical tests
support a forecasting design that combines stationary return targets, robust training cri
teria, market-index information, and explicit volatility modeling.
SOICT-HUST
Stock Price Analysis
40
Chapter 4
Forecasting Model Development
4.1 Problem Formulation
The forecasting task is formulated as a supervised learning problem in time series
analysis. Given a sequence of historical financial observations, the objective is to pre
dict future market movements over Let Xt denote the input feature vector constructed
from past observations, and yt+h denote the target variable at forecasting horizon h. The
learning objective can be expressed as
Xt =(xt−w,xt−w+1,…,xt−1) → yt+h
(4.1)
where w is the lookback window size and h is the forecasting horizon.
Since the original closing price series is generally non-stationary, directly forecasting
future prices may lead to unstable model performance. As stated in Hypothesis 1, the
log-return series exhibits stronger stationarity properties than the raw closing price series,
making it a more suitable target for supervised forecasting models. Based on the validation
of Hypothesis 1 through stationarity tests, we therefore define the forecasting target as
the future log return rather than the raw closing price.
The target variable is computed as
yt+h = rt+h = log(Pt+h) − log(Pt) = log Pt+h
Pt
,
(4.2)
where Pt denotes the closing price at time t. The model is trained to predict the
future log return rt+h. During inference, the predicted log return is transformed back to
the original price scale through the inverse logarithmic operation:
ˆ
P ∗t+h=Ptexp(ˆr∗t+h),
(4.3)
where ˆr∗t + h is the predicted log return and ˆP∗t + h is the reconstructed forecasted
price.
SOICT-HUST
Stock Price Analysis
41
4.2 Data Representation and Feature Engineering
In financial time series forecasting, particularly for the Vietnamese stock market (e.g.,
indices like VNINDEX, VN30 and individual equities such as FPT), raw price data
(OHLCV) is highly noisy (low signal-to-noise ratio) and exhibits non-stationarity. To
address this challenge, Feature Engineering serves as a critical phase to extract structural
signals, momentum, volatility, and market relationships, thereby providing high-value pre
dictive representations for machine learning (XGBoost) and deep learning (LSTM, GRU)
models.
The feature engineering workflow in this study is structured into three main compo
nents: Market Features, Stock Features, and the Data Processing & Pipeline.
4.2.1 Market Features
This group of features captures the macro state and general trend of the broad market
using the two major benchmarks of the Vietnamese stock market: VNINDEX and VN30.
Incorporating market features enables the predictive models to recognize the prevailing
market regime and adapt their individual stock forecasts accordingly.
• Market Momentum & Trend: The market trend is quantified using rates of
change (momentum) across different rolling lookback windows (w ∈ {5,10,20}
days):
Market_mom_wt = Pmarket,t −Pmarket,t−w
Pmarket,t−w
(4.4)
Additionally, a bullish market regime flag is defined by the crossover of short- and
long-term Exponential Moving Averages (EMA):
Market_bullt = I(EMA10(Pmarket)t−1 > EMA20(Pmarket)t−1)
The standardized trend strength of the market is expressed as:
Market_trend_strengtht = EMA10,t − EMA20,t
EMA20,t +ϵ
(4.5)
(4.6)
To measure the peak-to-trough decline over a rolling 20-day horizon, the market
drawdown is calculated as:
Market_drawdownt = Pmarket,t − maxi∈[0,19] Pmarket,t−i
maxi∈[0,19] Pmarket,t−i + ϵ
(4.7)
If the drawdown drops below-5%, the binary indicator Market_deep_drawdown is
activated to signal extreme systemic risk.
SOICT-HUST
Stock Price Analysis
42
• Market Volatility: Market price volatility is measured by the rolling standard
deviation of daily log returns, annualized by multiplying by √252:
σmarket,w,t = std(rmarket,[t−w+1,t]) ×
√
252
(4.8)
To flag periods of market panic or extreme volatility, a high volatility regime indi
cator is triggered when the short-term volatility exceeds its long-term average by
10%:
Market_is_high_vol_regimet = I(σmarket,20,t > 1.1 × MA20(σmarket,20)t) (4.9)
Furthermore, the asymmetry between upside and downside market volatility is cap
tured using Downside Volatility and Volatility Asymmetry:
downside_vol_20t = std({rmarket,τ | rmarket,τ < 0,τ ∈ [t − 19,t]}) ×
√
252 (4.10)
vol_asymmetryt = downside_vol_20t
σmarket,20,t + ϵ
(4.11)
• Stock-Market Beta & Relative Strength: The dynamic correlation between the
individual stock and the market is represented by a rolling 20-day Beta coefficient
(β20):
βw,t = Cov(rstock,rmarket)[t−w+1,t]
Var(rmarket)[t−w+1,t] + ϵ
(4.12)
The stability of the Beta coefficient is defined as its relative variation over the last
10 days:
Beta_stabilityt =
std(β20)[t−9,t]
|mean(β20)[t−9,t]| + ϵ
(4.13)
Relative Strength (RS) measures whether the stock is outperforming or underper
forming the broad market benchmark:
w−1
Relative_Strength_wt =
4.2.2 Stock Features
i=0
w−1
rstock,t−i −
i=0
rmarket,t−i
(4.14)
Stock-level features focus on describing the price action, volatility structure, and range
dynamics of the individual security. This group is categorized into several technical sub
classes:
• Trend & Moving Averages: We employ EMAs at short, medium, and long-term
horizons (w ∈ {5,10,20}). The relative distance between the closing price and the
SOICT-HUST
Stock Price Analysis
43
EMA provides a local overbought/oversold indicator:
price_ema_dist_wt = Closet − EMAw(Close)t
EMAw(Close)t + ϵ
The EMA slope measures trend acceleration:
ema_slope_w_st = EMAw(Close)t −EMAw(Close)t−s
s ×(Closet +ϵ)
(4.15)
(4.16)
• MACD(Moving Average Convergence Divergence): MACD calculations are
adapted to the specific forecast horizon h (for instance, when h = 5, the optimized
configuration is set to fast = 5,slow = 15,signal = 5):
MACDt = EMAfast(Close)t −EMAslow(Close)t
Signalt = EMAsignal(MACD)t
Histt = MACDt − Signalt
(4.17)
(4.18)
(4.19)
To allow models to learn divergence boundaries scale-invariantly, these components
are normalized by dividing by the closing price (macd_pct, hist_pct).
• Momentum & Oscillators:– RSI (Relative Strength Index): RSI is calculated over 7-day and 14-day win
dows. To transform threshold boundaries non-linearly (helping tree and linear
models partition the space), we apply a smoothed sigmoid function:
RSI_overbought_smootht =
RSI_oversold_smootht =
1
1 +e−0.2(RSI14,t−70)
1
1 +e0.2(RSI14,t−30)
(4.20)
(4.21)– Stochastic Oscillator: The %K line indicates the relative position of the closing
price within the High-Low range of the past 14 days:
Stoch_Kt = 100 ×
Closet − mini∈[0,13] Lowt−i
maxi∈[0,13] Hight−i − mini∈[0,13] Lowt−i + ϵ
(4.22)– Momentum Composite: To create a robust, noise-reduced momentum signal,
we combine multiple standardized Rate of Change (ROC) features:
mom_compositet = 0.3ROC5,t
σROC5,t
+ 0.4ROC10,t
σROC10,t
+ 0.3ROC20,t
σROC20,t
(4.23)
• Market Structure & Range Position: These features track the current price po
SOICT-HUST
Stock Price Analysis
44
sition relative to dynamic support and resistance levels across multiple timeframes:
range_position_wt =
Closet − mini∈[0,w−1] Closet−i
maxi∈[0,w−1] Closet−i − mini∈[0,w−1] Closet−i + ϵ
(4.24)
The breakout strength quantifies the percentage deviation of the current price be
yond the previous period’s high/low:
breakout_high_strength_wt = max Closet −High_prev_w
High_prev_w +ϵ ,0
(4.25)
• Volatility & Risk: This includes historical volatility, downside volatility, Average
True Range percentage (atr_pct) to measure price dispersion relative to market
price, and Volatility of Volatility (vol_of_vol_20) to capture the uncertainty of
risk.
• Return Lags & Distribution: Lags of past log returns model autocorrelation
structure, while rolling higher-order distribution moments such as Skewness, Kur
tosis, and Quantiles (quantile 25%, quantile 75%) capture the fat-tailed return dis
tributions typical of financial assets.
4.2.3 Data Processing & Pipeline
To ensure features are mathematically appropriate for model input, avoid look-ahead
bias, and improve convergence during training, the data pipeline implements the following
systematic steps:
1. Target Construction: The target variable (yt) is defined as the cumulative log
return over the forecast horizon h:
yt = ln Closet+h
Closet
(4.26)
Using log returns guarantees time-additivity and shifts the distribution of the target
variable closer to normality.
2. Safe Rolling & EWM Operators:Toprevent missing data gaps (due to holidays,
trading halts) from propagating NaN values across rolling computations, we deploy
specialized safe rolling operations (safe_rolling and safe_ewm) with a minimum
data parameter min_pct = 0.8. This requires at least 80% valid observations in the
window to return a value.
3. Leakage Validation: To eliminate look-ahead bias, we execute automated corre
SOICT-HUST
Stock Price Analysis
45
lation checks between features at time t and future targets yt+k (where k ≥ 1):
ρk = Corr(Xi,t,yt+k)
(4.27)
If any feature displays a future correlation |ρk| > 0.3 for k ≥ 1, the pipeline flags it
for investigation.
4. Warm-up Period Handling:Rollingindicators (especially 20-day or 252-day met
rics) require sufficient historical context for initialization. A warm-up period of 252
trading days (approx. one business year) is discarded at the start of the dataset
before model partitioning to remove initial NaN values.
5. Feature Selection & Collinearity Mitigation: Given an initial feature space
exceeding 300 engineered indicators, a three-stage feature selection pipeline was
employed to identify a compact, non-redundant, and predictive feature subset for
each forecast horizon h ∈ {1,5,10}.
Stage 1– Importance Ranking via Gradient Boosting.AnXGBoostregressor
was trained on the full feature set using a chronological (non-shuffled) train/test split
to preserve temporal causality and prevent look-ahead bias. Feature importance was
computed using the gain-based criterion, and the top 20–30 features were retained
as candidates for each horizon.
Stage 2– Pairwise Collinearity Filtering. To mitigate multicollinearity– which
is known to destabilize both tree-based importance estimates and the weight estima
tion of downstream models such as MLP and linear baselines– a Pearson correlation
matrix was computed over the candidate feature subset. For any pair of features ex
ceeding an absolute correlation threshold of |r| > 0.85, the feature with the lower
XGBoost importance score was discarded, retaining the more predictive represen
tative of each correlated cluster.
Stage 3– SHAP-based Validation. To verify the robustness of the retained
feature set and detect residual cases where a single feature within a correlated group
disproportionately dominates the importance signal (a known artifact of tree-based
gain importance), the model was retrained on the filtered subset and SHAP (SHapley
Additive exPlanations) values were computed on the held-out test set. The mean
absolute SHAP value was used as a second, model-agnostic ranking criterion, and
the consistency between gain-based importance and SHAP rankings was used as a
stability check.
This procedure was applied independently to each forecast horizon, yielding horizon
specific feature subsets of 16–18 features after collinearity filtering (detailed in Ta
ble 4.1). To support a unified modeling pipeline across all three horizons, a con
SOICT-HUST
Stock Price Analysis
46
(a) h = 1 day
(b) h = 5 days
(c) h = 10 days
Figure 4.1: SHAP summary plots of the selected features for forecasting horizons of 1,
5, and 10 trading days. Features are ranked by mean absolute SHAP value, indicating
their overall contribution to model predictions. The variation in feature rankings across
horizons highlights the horizon-dependent nature of predictive signals in financial markets.
solidated set of 20 features was further derived by aggregating per-horizon SHAP
rankings– features consistently ranked highly across multiple horizons were priori
tized, with horizon-specific high-impact features added to complete the final set.
Based on the SHAP rankings obtained across the three forecasting horizons, a con
solidated set of 20 features was selected for the final modeling pipeline. Features that
consistently exhibited high SHAP importance across multiple horizons were priori
tized, while a small number of horizon-specific predictors were retained to preserve
predictive information unique to particular forecasting windows.
The selected features collectively capture trend, momentum, volatility, distributional
characteristics, price positioning, and broader market regime information, provid
ing a compact yet informative representation of market dynamics for downstream
forecasting models.
6. Data Splitting & StandardScaler: Data is divided chronologically to preserve
the temporal properties of financial series and prevent leakages from future to past.
The dataset split ratios are: Train (80%), Validation (10%), and Test (10%).
Z-score normalization is applied to each feature:
Xscaled = X −µtrain
σtrain
(4.28)
The scaling statistics (µtrain and σtrain) are computed exclusively on the Train set,
and then applied to transform the Train, Validation, and Test sets. This guarantees
no future information from Validation/Test leaks into model training.
SOICT-HUST
StockPriceAnalysis 47
Table4.1:SummaryofSelectedEngineeredFeatures
FeatureName FeatureGroup MathematicalDefinition Predictive Pur
pose
close Price Currentclosingpriceofthestock Baselineassetvalua
tion.
vol_10 StockVolatility Annualized10-dayrollingreturn
standarddeviation
Captures short-term
volatility.
ema_20 StockTrend 20-dayExponentialMovingAv
erageofprice
Identifies medium
termtrend.
RSI_price_corr_20 StockMomentum 20-day Pearson correlation be
tweenRSI-14andprice
Measures diver
gence/alignment of
momentum.
return_q25_20 ReturnDistribution Rolling25thpercentileofreturns
over20days
Captures downside
tailrisk.
return_q75_20 ReturnDistribution Rolling75thpercentileofreturns
over20days
Captures upside
growthpotential.
return_max_15 ReturnDistribution Maximumreturnover a rolling
15-daywindow
Capturespositiveex
tremeshocks.
return_std_10 StockVolatility 10-dayrollingstandarddeviation
of logreturns
Quantifies return
dispersion.
return_mom_10_20 StockMomentum Mean(R)10−Mean(R)20 Shorter- vs longer
term momentum
crossover.
dist_to_high_3 RangeStructure (Close−Max3,t−1)/Max3,t−1 Distanceto3-daylo
calresistance.
dist_to_high_5 RangeStructure (Close−Max5,t−1)/Max5,t−1 Distanceto5-daylo
calresistance.
downside_vol_10 StockVolatility 10-dayrollingsemi-standardde
viationofnegativereturns
Focusesondownside
risk.
downside_vol_20 StockVolatility 20-dayrollingsemi-standardde
viationofnegativereturns
Focuses on interme
diatedownsiderisk.
vn30_beta_20 MarketRelation 20-day rolling Beta coefficient
relativetoVN30Index
Captures systemic
riskrelativetolarge
caps.
vnindex_trend_strength MarketTrend TrendstrengthofVNINDEXin
dex((EMA10−EMA20)/EMA20)
Market trend envi
ronment indicator.
vnindex_vol_asymmetry MarketVolatility Downside-to-totalvolatilityratio
ofVNINDEXindex
Marketriskasymme
tryindicator.
vn30_vol_10 MarketVolatility 10-dayrollingvolatilityofVN30
Index(annualized)
Volatility in the
blue-chip market
segment.
vn30_is_high_vol_regime MarketVolatility Flag if VN30 volatility exceeds
long-termaverageby10%
Identifies market
turbulenceregimes.
breakout_low_strength_10 RangeStructure max((Min10,t−1 −
Close)/Min10,t−1,0)
Strength of support
breakouts over 10
days.
return_cumsum_20 ReturnDistribution Cumulative log returnover the
past20days
Intermediate-term
priceperformance.
return_min_10 ReturnDistribution Minimumreturn over a rolling
10-daywindow
Captures negative
extremeshocks.
SOICT-HUST
Stock Price Analysis
48
4.3 Model Selection
4.3.1 Summary of Selected Features
Table 4.1 summarizes the selected core features utilized in model training.
To establish a robust predictive framework for financial time series forecasting, we eval
uate and compare models across three paradigms: classical statistical models, ensemble
based machine learning approaches, and deep learning architectures. This diverse selection
allows us to compare parametric baselines against non-parametric estimators capable of
modeling complex non-linear patterns.
4.3.2 Classical Models
Classical econometric models serve as parametric baselines. They rely on formal sta
tistical assumptions about the underlying data generating process, specifically modeling
linear autocorrelation and time-varying variance structures.
• ARIMA (Autoregressive Integrated Moving Average): ARIMA is the stan
dard parametric model for forecasting stationary univariate time series. For a dif
ferenced series Yt = (1 − B)dXt, where B is the backshift operator (BXt = Xt−1)
and d is the order of integration, the model ARIMA(p,d,q) is formalized as:
ϕ(B)Yt = c+θ(B)ϵt
(4.29)
where ϕ(B) = 1− p
i=1
ϕiBi represents the autoregressive (AR) polynomial of order
p, θ(B) = 1+ q
j=1
θjBj represents the moving average (MA) polynomial of order
q, and ϵt ∼ WN(0,σ2) is a white noise process representing random innovations. In
this study, ARIMA is fit directly on the stationary log returns (d = 0), leveraging
past returns to capture short-term linear dependencies.
• GARCH (Generalized Autoregressive Conditional Heteroskedasticity):
Financial return series commonly exhibit volatility clustering, where high-volatility
periods group together, violating the homoskedasticity assumption of standard re
gression. The GARCH(p,q) model addresses this by modeling the conditional vari
ance σ2
t as an autoregressive process of past squared residuals and past conditional
variances:
rt = µt +ϵt, ϵt = σtzt, zt ∼ i.i.d.N(0,1)
q
σ2
t = ω +
i=1
p
αiϵ2
t−i +
j=1
βjσ2
t−j
(4.30)
(4.31)
SOICT-HUST
Stock Price Analysis
49
where ω > 0, αi ≥ 0, and βj ≥ 0 are parameters to ensure positive conditional
variance, and
αi+ βj <1 guarantees covariance stationarity. GARCH captures
the time-varying volatility dynamics and fat-tailed distribution profiles inherent in
asset returns.
4.3.3 Machine Learning Models
Machine learning models are non-parametric, relaxing linear assumptions to capture
complex non-linear feature interactions and high-dimensional relationships.
• RandomForest (RF):RandomForest is an ensemble bootstrap aggregation (bag
ging) model composed of a collection of randomized decision trees. For regression,
it reduces variance without increasing bias by averaging the predictions of M inde
pendent trees:
M
ˆ
fRF(x) = 1
M
m=1
Tm(x;Θm)
(4.32)
where Θm represents the parameter set (split variables, split points, and terminal
node values) of the m-th tree trained on a bootstrap sample of the training data.
At each node split, a random subset of features is evaluated, which decorrelates the
trees and increases robustness against overfitting.
• XGBoost (Extreme Gradient Boosting): XGBoost is an efficient, scalable im
plementation of gradient-boosted decision trees. Unlike bagging, it trains trees se
quentially to minimize a regularized objective function:
n
Obj(t) =
i=1
l(yi, ˆy(t−1)
i
+ft(xi)) + Ω(ft)
(4.33)
where l is the loss function measuring the deviation of the prediction at step t
(ˆy(t−1)
i
+ft(xi)) from the target yi. The regularization term Ω(ft) penalizes model
complexity to prevent overfitting:
Ω(ft) = γT + 1
2λ
T
j=1
w2
j
(4.34)
where T is the number of terminal leaves and w represents leaf weights. XGBoost
handles collinearity and tabular interactions extremely well, using a second-order
Taylor expansion to optimize the objective function.
SOICT-HUST
Stock Price Analysis
50
4.3.4 Deep Learning Models
Deep learning models leverage neural network architectures to automatically learn hi
erarchical representations from sequential inputs without requiring manual feature align
ment over time.
• LSTM (Long Short-Term Memory): LSTM is a gated Recurrent Neural Net
work (RNN) designed to mitigate the vanishing gradient problem when capturing
long-term temporal dependencies in sequence data. It introduces a cell state Ct and
three gating mechanisms (forget gate ft, input gate it, and output gate ot):
ft = σ(Wf ·[ht−1,xt] + bf)
it = σ(Wi ·[ht−1,xt] + bi)
˜
Ct = tanh(WC ·[ht−1,xt] + bC)
Ct = ft ⊙Ct−1 +it ⊙ ˜Ct
ot = σ(Wo ·[ht−1,xt] + bo)
ht = ot ⊙tanh(Ct)
(4.35)
(4.36)
(4.37)
(4.38)
(4.39)
(4.40)
where σ is the sigmoid activation, ⊙ represents element-wise multiplication, and
ht is the hidden state. In our implementation, the final hidden state of the LSTM
sequence h(−1)
n
is extracted and fed into a fully connected prediction head containing
a Layer Normalization step, a ReLU activation, and a Dropout layer to output the
predicted log return.
• GRU(Gated Recurrent Unit): GRU is a streamlined variant of the LSTM that
merges the cell state and hidden state, and combines the forget and input gates into
a single update gate zt, alongside a reset gate rt:
zt = σ(Wz ·[ht−1,xt] + bz)
rt = σ(Wr ·[ht−1,xt] + br)
˜
ht = tanh(W ·[rt ⊙ht−1,xt] +bh)
ht = (1 −zt)⊙ht−1 +zt ⊙˜ ht
(4.41)
(4.42)
(4.43)
(4.44)
With fewer parameters than LSTM, GRU is less prone to overfitting on smaller
financial datasets and trains faster while achieving comparable representation ca
pacity. Similar to the LSTM, the last hidden state h(−1)
n
projection head to yield the final forecast.
is processed through a
SOICT-HUST
Stock Price Analysis
51
4.4 Hyperparameter Optimization
Due to the extremely low signal-to-noise ratio in financial data, predictive models
are highly sensitive to hyperparameter configurations. Suboptimal settings can lead to
severe overfitting or underfitting. To address this, we implement a systematic hyperpa
rameter optimization (HPO) framework using **Optuna**, a state-of-the-art Bayesian
optimization framework.
4.4.1 Optimization Methodology
Our framework utilizes the Tree-structured Parzen Estimator (TPE) sampler as its
core search engine. TPE is a sequential model-based optimization (SMBO) algorithm
that models the probability of hyperparameters belonging to two distinct distributions:
one yielding objective values above a certain threshold (good configurations), and one
yielding values below (poor configurations). The sampler dynamically updates these prob
ability distributions after each trial, prioritizing regions of the search space that maximize
expected improvement.
As established by Hypothesis 3, the Mean Absolute Error (MAE) computed on recon
structed closing prices provides a more appropriate evaluation criterion than RMSE for
practical forecasting applications, as it is less sensitive to extreme prediction errors and
better reflects the typical magnitude of forecast deviations. Consequently, the hyperpa
rameter optimization process is designed to minimize the validation MAE of the recon
structed closing prices rather than metrics computed directly on log returns or RMSE
based objectives.
The objective metric evaluated for each trial is the reconstructed validation MAE:
MAEreconstructed = 1
N
N
i=1
|Pi,actual − Pi,pred|
(4.45)
By evaluating error directly on price levels, the optimization process accounts for com
pounding effects and market scaling.
4.4.2 Hyperparameter Search Spaces
For each model family, we define a structured search space targeting parameters that
control learning dynamics, model capacity, and regularization.
• Recurrent Architectures (LSTM & GRU): The recurrent models’ capacity
is determined by the size and depth of their hidden layers, while generalization is
controlled by dropout rates and learning steps.
SOICT-HUST
Stock Price Analysis
52– Hidden Size (hidden_size): Categorical variable ∈ {16,32,64,128}, governing
the size of the hidden state vectors.– Number of Layers (num_layers): Integer range [1,3], controlling the stacking
depth of the recurrent cells.– Dropout (dropout): Range [0.0,0.5] (categorical for LSTM, uniform float for
GRU), regularizing connections between recurrent layers.– Learning Rate (lr): Log-uniform range [10−5,10−2], adjusting the optimization
steps of the Adam optimizer.– Batch Size (batch_size): Categorical variable ∈ {16,32,64,128}, adjusting gra
dient estimation resolution.
• Gradient Boosting (XGBoost): XGBoost tuning focuses on finding a balance
between individual tree complexity, the total number of trees, and stochastic regu
larization parameters.– Estimators (n_estimators): Integer range [50,500], controlling the number of
boosting iterations.– Max Depth (max_depth): Integer range [3,10], limiting the maximum depth of
individual decision trees.– Learning Rate (learning_rate): Log-uniform float range [0.01,0.3], scaling the
contribution of each subsequent tree (shrinkage).– Subsample (subsample): Uniform float range [0.5,1.0], representing the fraction
of training instances sampled stochastically per tree.– Column Sample (colsample_bytree): Uniform float range [0.5,1.0], representing
the fraction of features sampled stochastically per tree.– Minimum Child Weight (min_child_weight): Integer range [1,10], representing
the minimum sum of instance weight needed in a child node to allow further
partitioning.
4.4.3 Summary of Search Space Parameters
Table 4.2 provides a comprehensive overview of the search parameters, their data
types, search intervals, and their primary role in model regularisation.
SOICT-HUST
StockPriceAnalysis 53
Table4.2:HyperparameterSearchSpaceSpecifications
Model Parameter Type Search Space /
Range
Theoretical Pur
pose
LSTM/GRU
hidden_size Categorical {16,32,64,128} Controlsmodel rep
resentationcapacity.
num_layers Integer [1,3] Adjustsmodeldepth
to learnhierarchical
temporal features.
dropout Float/Cat [0.0,0.5] Prevents co
adaptation of
weights in recur
rentunits.
learning_rate Log-Float [10−5,10−2] Governs themagni
tudeof gradientup
dates.
batch_size Categorical {16,32,64,128} Controls gradient
variance during
training.
XGBoost
n_estimators Integer [50,500] Controls the ensem
ble capacity (boost
ingrounds).
max_depth Integer [3,10] Restricts the inter
actiondepthof fea
tures.
learning_rate Log-Float [0.01,0.3] Restricts step up
datestopreventvari
anceinflation.
subsample Float [0.5,1.0] Introduces instance
bagging to prevent
overfitting.
colsample_bytree Float [0.5,1.0] Decorrelatestreesby
columnbagging.
min_child_weight Integer [1,10] Restricts tree split
tingbasedonsample
coverage.
SOICT-HUST
Stock Price Analysis
54
Chapter 5
Model Evaluation and Conclusion
5.1 Forecasting Performance Evaluation
This chapter evaluates the forecasting performance of the proposed forecasting frame
work using three machine learning and deep learning models: XGBoost, LSTM, and GRU.
The evaluation is conducted on the FPT stock dataset under multiple forecasting hori
zons. In addition to traditional forecasting error metrics, statistical significance tests and
confidence interval analysis are performed to assess the reliability and robustness of the
obtained results.
5.1.1 Experimental Setup
Following the feature engineering and feature selection stages described in Chapter 4,
the final feature set is used to construct supervised learning samples through a sliding
window framework. Given a lookback window of size w, the forecasting task can be for
mulated as:
Xt =[yt−w,…,yt−1], yt = f(Xt)
Three forecasting horizons are considered:
• One-day ahead forecasting (h = 1)
• Five-day ahead forecasting (h = 5)
• Ten-day ahead forecasting (h = 10)
(5.1)
For XGBoost, LSTM, and GRU models, Min-Max normalization is applied to the
training data:
x′ = x−xmin
xmax −xmin
(5.2)
SOICT-HUST
Stock Price Analysis
55
All predicted values are subsequently transformed back to the original price scale
before evaluation.
5.1.2 Evaluation Metrics
The forecasting models are evaluated using four commonly adopted metrics.
Mean Absolute Error (MAE)
MAE = 1
n
n
t=1
|yt − ˆyt|
Root Mean Squared Error (RMSE)
n
RMSE =
1
n
t=1
(yt − ˆyt)2
Mean Absolute Percentage Error (MAPE)
MAPE = 1
n
n
t=1
yt − ˆyt
yt
Directional Accuracy (DA)
×100
DA= 1
n
n
t=1
I(sign(ˆyt − yt−1) = sign(yt − yt−1))
(5.3)
(5.4)
(5.5)
(5.6)
RMSE and MAE evaluate forecasting accuracy, MAPE provides a scale-independent
error measure, while Directional Accuracy evaluates the ability of the model to predict
market direction.
5.1.3 Forecasting Results
Table 5.1 presents the forecasting performance of the three models across different
forecasting horizons.
“‘latex
SOICT-HUST
StockPriceAnalysis 56
Table5.1:ForecastingPerformanceAcrossDifferentHorizons
Model Horizon RMSE MAPE MAE DA
LSTM
1 2.0990(1.5686,2.6958) 1.3598%(1.0164%,1.8095%) 1.4596(1.1493,1.8388) 47.09%(47.09%,56.08%)
5 4.7720(3.833,5.954) 3.458%(2.67%,4.51%) 3.7225(3.004,4.720) 53.19%(47.9%,61.2%)
10 6.9531(5.381,8.907) 5.093%(3.87%,6.92%) 5.5050(4.317,7.237) 51.60%(47.9%,60.7%)
GRU
1 2.0326(1.5886,2.5523) 1.3237%(1.0187%,1.7259%) 1.4322(1.1508,1.8087) 48.68%(48.15%,58.20%)
5 4.2769(3.296,5.450) 2.982%(2.29%,4.04%) 3.2312(2.557,4.230) 49.47%(46.3%,58.0%)
10 5.3957(4.050,6.907) 3.723%(2.80%,5.07%) 4.0676(3.140,5.354) 51.60%(47.3%,61.2%)
XGBoost
1 1.9226(1.5159,2.3899) 1.2284%(0.9502%,1.5883%) 1.3332(1.0731,1.6618) 47.47%(46.97%,57.58%)
5 4.3004(3.263,5.462) 2.931%(2.14%,4.01%) 3.1623(2.389,4.192) 50.53%(46.8%,59.6%)
10 5.9617(4.452,7.607) 3.8943%(2.7273%,5.5140%) 4.3581(3.195,5.973) 52.66%(47.9%,62.8%)
ARIMA
1 1.9371(1.496,2.446) 1.217%(0.94%,1.59%) 1.3222(1.065,1.662) 46.03%(46.0%,55.0%)
5 4.1738(3.188,5.334) 2.866%(2.15%,3.90%) 3.1064(2.415,4.096) 50.00%(46.8%,58.5%)
10 5.9237(4.388,7.655) 4.100%(2.96%,5.77%) 4.4401(3.314,6.005) 52.66%(48.4%,62.8%)
GARCH
1 1.9384(1.497,2.448) 1.218%(0.94%,1.59%) 1.3235(1.066,1.662) 46.03%(46.0%,55.0%)
5 4.2310(3.240,5.400) 2.930%(2.22%,3.99%) 3.1755(2.492,4.175) 48.40%(45.2%,56.9%)
10 5.9454(4.401,7.689) 4.118%(2.99%,5.82%) 4.4598(3.348,6.044) 50.53%(46.3%,60.1%)
“‘
SeveralobservationscanbedrawnfromTable5.1.
Theforecastingresultsdemonstratethatmodelperformancevariesacrossforecasting
horizons.Forthe1-dayhorizon,XGBoostachievedthelowestRMSE,whileARIMAand
GARCHproducedthelowestMAEandMAPEvalues, indicatingstrongshort-termfore
castingcapability.Atthe5-dayhorizon,ARIMAobtainedthelowestMAEandMAPE,
whereasGRUachievedthelowestRMSE, suggestingthatrecurrentneuralnetworksare
more effectiveat capturingmedium-termtemporal dynamics.For the10-dayhorizon,
GRUconsistentlydeliveredthelowestRMSE,MAE,andMAPEvalues,highlightingits
superiorabilitytomodel longer-termdependenciesinfinancial timeseries.
Regardingdirectional accuracy, thedifferencesamongmodelswere relativelysmall,
withmostvaluesremainingcloseto50
5.2 StatisticalReliabilityofForecasts
Whileforecastingmetricsprovideusefulinformationregardingpredictiveperformance,
statistical testingisrequiredtodeterminewhetherobserveddifferencesbetweenmodels
aremeaningful.
5.2.1 BootstrapConfidenceIntervals
Toevaluatetherobustnessofforecastingperformance,bootstrapresamplingisapplied
tothepredictionresiduals:
et=yt−ˆyt (5.7)
The95
SOICT-HUST
StockPriceAnalysis 57
CI95%=[P2.5,P97.5] (5.8)
Thebootstrapconfidenceintervalsreveal substantialoverlapbetweenGRUandXG
Boostacrossmostforecastinghorizons.Therefore,althoughonemodelmayexhibitalower
pointestimateofforecastingerror,thetrueperformancedifferencemayberelativelysmall
onceestimationuncertaintyisconsidered.
5.2.2 Diebold–MarianoTest
Todeterminewhetherforecastingperformancedifferencesarestatisticallysignificant,
pairwiseDiebold–Mariano(DM)testsareconducted.
Thenullhypothesisisdefinedas:
H0 :E=0 (5.9)
wheredtdenotesthelossdifferentialbetweentwocompetingmodels.
Tables5.2and5.3summarizetheresultingp-values.
Table5.2:Diebold–MarianoTestResults(RMSEandMAE)
Horizon ModelPair RMSEp-value MAEp-value Significant
1 GRUvsLSTM 0.0063 0.0290 Yes
1 GRUvsXGBoost 0.0085 0.0156 Yes
1 LSTMvsXGBoost 0.6768 0.3372 No
5 GRUvsLSTM 0.0498 0.0484 Yes
5 GRUvsXGBoost 0.4711 0.8243 No
5 LSTMvsXGBoost 0.0360 0.0776 Partial
10 GRUvsLSTM 0.0140 0.0344 Yes
10 GRUvsXGBoost 0.3730 0.1296 No
10 LSTMvsXGBoost 0.0067 0.0498 Yes
Table5.3:Diebold–MarianoTestResultsAgainstStatisticalModels
Horizon ModelPair RMSEp-value MAEp-value Significant
1 GRUvsARIMA 0.0429 0.0106 Yes
1 GRUvsGARCH 0.0441 0.0116 Yes
1 ARIMAvsGARCH 0.1092 0.1204 No
5 GRUvsARIMA 0.0246 0.0022 Yes
5 GRUvsGARCH 0.0285 0.0051 Yes
5 ARIMAvsGARCH 0.0271 0.0021 Yes
10 GRUvsARIMA 0.2123 0.2807 No
10 GRUvsGARCH 0.2049 0.2676 No
10 ARIMAvsGARCH 0.1050 0.2987 No
TheDiebold–Marianotestwasconductedtodeterminewhethertheobserveddiffer
ences inforecastingerrorswerestatisticallysignificant.The results indicate thatGRU
SOICT-HUST
Stock Price Analysis
58
significantly outperformed LSTM across all forecasting horizons under both RMSE and
MAE loss functions, demonstrating the effectiveness of the GRU architecture for this
forecasting task.
For the 1-day horizon, significant differences were identified between GRU and the
competing models, while no significant differences were observed among XGBoost, ARIMA,
and GARCH. This suggests that these models achieved comparable short-term forecast
ing performance despite small differences in error metrics. At the 5-day horizon, GRU
remained significantly different from ARIMA and GARCH, whereas the performance gap
between GRU and XGBoost was not statistically significant. For the 10-day horizon,
although GRU achieved the lowest forecasting errors, its performance was statistically
comparable to XGBoost, ARIMA, and GARCH according to the Diebold–Mariano test.
These findings highlight the importance of complementing traditional error metrics
with statistical significance testing. Lower forecasting errors do not necessarily imply
statistically superior predictive performance, particularly when the differences between
competing models are relatively small.
5.3 Conclusion
This study proposed a statistically grounded framework for Vietnamese stock price
analysis and forecasting, using FPT as the main case study and VNINDEX and VN30
as market references. The results show that FPT closing prices are non-stationary, while
log returns are stationary. FPT returns also exhibit heavy tails and strong volatility
clustering, which supports the use of return-based modeling, robust evaluation metrics,
and volatility-aware features.
The project combines statistical hypothesis testing, signal processing, feature engineer
ing, and predictive modeling. ARIMA/SARIMAX and GARCH-type models are used as
classical baselines, while XGBoost, MLP, LSTM, and GRU are used to capture nonlinear
and sequential patterns. Model performance is evaluated through error metrics, directional
accuracy, statistical comparison tests, confidence intervals, and multi-horizon forecasting.
Overall, the study shows that stock forecasting should be treated not only as a pre
diction problem, but also as an experimental design problem. Statistical testing helps
determine appropriate transformations and assumptions, while machine learning provides
flexible tools for prediction.
5.4 Future Work
Future work can extend this study in several directions. First, the analysis should be
expanded from FPT to more Vietnamese large-cap stocks across different sectors. Sec
ond, volatility-specific models such as GARCH, EGARCH, or ARIMA–GARCH hybrids
SOICT-HUST
Stock Price Analysis
59
should be further developed. Third, additional exogenous variables such as interest rates,
exchange rates, foreign trading flows, macroeconomic indicators, and news sentiment can
be added. Fourth, trading-oriented metrics such as Sharpe ratio, maximum drawdown, cu
mulative return, turnover, and transaction-cost-adjusted performance should be included.
Finally, model interpretability methods such as feature importance or SHAP values can
be used to explain which factors drive model predictions.
SOICT-HUST
Stock Price Analysis
60
Bibliography
[1] G. E. P. Box, G. M. Jenkins, G. C. Reinsel, and G. M. Ljung, Time Series Analysis:
Forecasting and Control. John Wiley & Sons, 5th ed., 2015.
[2] D. A. Dickey and W. A. Fuller, “Distribution of the estimators for autoregressive
time series with a unit root,” Journal of the American Statistical Association, vol. 74,
no. 366a, pp. 427–431, 1979.
[3] R. F. Engle, “Autoregressive conditional heteroscedasticity with estimates of the vari
ance of united kingdom inflation,” Econometrica: Journal of the Econometric Society,
vol. 50, no. 4, pp. 987–1007, 1982.
[4] T. Bollerslev, “Generalized autoregressive conditional heteroskedasticity,” Journal of
Econometrics, vol. 31, no. 3, pp. 307–327, 1986.
[5] E. F. Fama, “Efficient capital markets: A review of theory and empirical work,” The
Journal of Finance, vol. 25, no. 2, pp. 383–417, 1970.
[6] W. F. Sharpe, “Capital asset prices: A theory of market equilibrium under conditions
of risk,” The Journal of Finance, vol. 19, no. 3, pp. 425–442, 1964.
[7] T. Chen and C. Guestrin, “Xgboost: A scalable tree boosting system,” in Proceedings
of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and
Data Mining, pp. 785–794, ACM, 2016.
[8] S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation,
vol. 9, no. 8, pp. 1735–1780, 1997.
[9] K. Cho, B. van Merri¨enboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk,
and Y. Bengio, “Learning phrase representations using rnn encoder–decoder for sta
tistical machine translation,” in Proceedings of the 2014 Conference on Empirical
Methods in Natural Language Processing (EMNLP), pp. 1724–1734, 2014.
[10] F. X. Diebold and R. S. Mariano, “Comparing predictive accuracy,” Journal of Busi
ness & Economic Statistics, vol. 13, no. 3, pp. 253–263, 1995.
SOICT-HUST
Stock Price Analysis
61
[11] P. Welch, “The use of fast fourier transform for the estimation of power spectra: a
method based on time averaging over short, modified periodograms,” IEEE Trans
actions on audio and electroacoustics, vol. 15, no. 2, pp. 70–73, 1967.
[12] D. N. Politis and J. P. Romano, “The stationary bootstrap,” Journal of the American
Statistical Association, vol. 89, no. 428, pp. 1303–1313, 1994.
[13] M. Loève, Fonctions aléatoires de second ordre. Paris: Hermann, 1948.
[14] N. E. Huang, Z. Shen, S. R. Long, M. C. Wu, H. H. Shih, Q. Zheng, N.-C. Yen,
C. C. Chao, and H. H. Liu, “The empirical mode decomposition and the hilbert
spectrum for nonlinear and non-stationary time series analysis,” Proceedings of the
Royal Society of London. Series A: Mathematical, Physical and Engineering Sciences,
vol. 454, no. 1971, pp. 903–995, 1998.
[15] A. Grossmann and J. Morlet, “Decomposition of hardy functions into wavelets of
constant shape,” in SIAM Journal on Mathematical Analysis, vol. 15, pp. 723–736,
SIAM, 1984.
[16] C. Torrence and G. P. Compo, “A practical guide to wavelet analysis,” Bulletin of
the American Meteorological Society, vol. 79, no. 1, pp. 61–78, 1998.
[17] G. E. P. Box and G. M. Jenkins, “Time series analysis: Forecasting and control,”
Holden-Day, 1976.
[18] T. Vu, “Vnstock: Python vietnamese stock market data downloader.” https://
github.com/thinh-vu/vnstock, 2024.
[19] K. Karhunen, “ ¨ Uber lineare methoden in der wahrscheinlichkeitsrechnung,” Annales
Academiae Scientiarum Fennicae. Series A. I. Mathematica-Physica, vol. 37, pp. 1
79, 1947.
SOICT-HUST