A COMPARATIVE STUDY OF MACHINE LEARNING METHODS FOR PORTFOLIO MANAGEMENT IN INDIAN EQUITY MARKETS
Full text
NANYANG TECHNOLOGICAL UNIVERSITY SCHOOL OF SOCIAL SCIENCES A COMPARATIVE STUDY OF MACHINE LEARNING METHODS FOR PORTFOLIO MANAGEMENT IN INDIAN EQUITY MARKETS Submitted by: Behl Arav U2231911H Garg Pranav U2231021L Saraf Krish U2231741H A Graduation Project submitted to the School of Social Sciences, Nanyang Technological University in partial fulfillment of the requirements for the Degree of Bachelor of Arts in Economics Academic Year: 2025/2026
FINAL REPORT Abstract Despite widespread applications of machine learning techniques to financial data, implementations in the context of the Indian market utilising various data sources are scarce. This project aims to compare the effectiveness of various machine learning and reinforcement learning techniques in portfolio management for equities listed on the National Stock Exchange (NSE), India. We make use of multimodal data in the form of historical stock prices, technical indicators, financial data, news headlines and social media posts from the period of June 2024 to June 2025. We investigate four methodologies using: Ridge regression, a Large Language Model (LLM), a Long Short Term Memory (LSTM) neural network, and a Reinforcement Learning model with the Proximal Policy Optimisation (PPO) agent. We find that all the aforementioned models provide higher returns than benchmark traditional portfolio optimisation techniques. However, we also note that the PPO model provides only marginal gains compared to Ridge and LSTM, while being much more computationally intensive. Our project provides a set-up for future practical implementations of these machine learning techniques on a larger scale. I. Introduction Financial data from stock markets is noisy and complex. Historical prices, financial fundamentals, technical indicators, and sentiment from news and social media all contain valuable signals, but these signals are difficult to extract using traditional methods. Machine learning has therefore been extensively applied to portfolio management, with techniques such as neural networks and deep reinforcement learning showing promise in developed markets. Using deep learning methods, as discussed in the methodology section, can uncover latent complex interactions between variables in our dataset, which may be jointly associated with stock prices and movements. India presents a unique and compelling context for automated portfolio management. The number of retail investors in Indian equity markets has grown rapidly, with the number of demat accounts increasing from 40 million in 2020 to over 150 million by 2024 (IBEF, 2024; Business India, 2024). As of March 2025, the National Stock Exchange (NSE) has 113 million registered investors (News4Masses, 2025), with retail investors now contributing over 45% of daily cash market turnover (Outlook Business, 2025). Young investors under 30 have grown from 29% of all investors in FY19 to 48% in FY23 (IBEF, 2024). This rapid expansion in equity investments among citizens in India can
lead to a strong demand for practical, automated portfolio management solutions that can balance risk and return. However, advanced machine learning applications remain limited in Indian markets compared to developed markets. India's market also exhibits unique characteristics as an emerging economy: higher volatility compared to developed markets, distinct regulatory frameworks under SEBI, and different investor behavior patterns. With India's market capitalization reaching $5.18 trillion in 2024, making it the world's fifth-largest market (ICICI Direct, 2025), understanding which machine learning techniques work effectively in this context is valuable for both researchers and practitioners. We study a variety of machine learning techniques, including regularized linear models, large language models, neural networks and reinforcement learning, aiming to meet the practical need for comparisons of their effectiveness in the Indian market. This project addresses the research question: how can machine learning be used to create a portfolio management system using multimodal data including historical price data, financial indicators, technical indicators, and news and social media sentiment, in the Indian equities market? We implement and compare multiple approaches across 50 equities representative of the NSE, and comprising the Nifty 50 index, evaluating their performance in terms of returns, volatility and Sharpe ratio to provide actionable insights for both researchers and investors. II. Literature Review The extant literature on portfolio management in equity markets discusses a variety of methodologies, such as traditional quantitative techniques, neural networks for the price predictions of individual stocks, machine learning and reinforcement learning methods to optimise portfolio returns, and LLMbased trading decisions for individual stocks. We summarise each of these methods in this section, leading to the identification of our research gap. Traditional Quantitative Techniques: Traditional portfolio management techniques include the BlackLitterman model (Black and Litterman, 1992), mean-variance optimisation (Markowitz, 1952), momentum-based strategy (Jegadeesh and Titman, 1993), among others. We discuss these techniques in more detail in the methodology section. Meher & Mishra (2024) apply the Black-Litterman model to large-cap equities from the Nifty 50 index of the National Stock Exchange (NSE), India, Bitcoin, a Gold ETF and a natural resources focused mutual fund, comparing it to Monte-Carlo simulations (Metropolis and Ulam, 1949) of various combinations of these assets to form other hypothetical
portfolios. They report that the best performing portfolio in the Monte-Carlo approach provides higher returns at a higher volatility than the Black-Litterman model. Pelluri (2025) investigates applying momentum-based strategies to midcap equities listed on the NSE, reporting cumulative returns far exceeding the Nifty 50 index. These techniques often require estimates of the expectations for future returns. Predictions of changes in stock prices using machine learning techniques could be especially useful in this context. Neural Networks for Price Prediction: Long Short Term Memory (LSTM) is the most widely used architecture for price and return predictions for individual stocks, due to its ability to retain the most relevant historical data at each timestep. Gu et al. (2020) demonstrate that LSTM networks outperform linear and tree-based models, in terms of both prediction accuracy and portfolio return, using extensive price, technical, financial and macroeconomic data from 1957 to 2016 for stocks listed in the NYSE, AMEX and NASDAQ. Chaudhary (2025) utilises news sentiment analysis using VADER (Valence Aware Dictionary for Sentiment Reasoning) (Hutto and Gilbert, 2014), and historical price data, with an LSTM model to predict returns for Google, Apple, Amazon, and Microsoft, achieving high accuracy with MAPE at about 3.1%. Fjellström (2022) uses an ensemble of LSTM networks and historical price data to predict whether the price of a stock on the OMX30 index of the Nasdaq Stockholm exchange is higher or lower than its median. A portfolio is then constructed by assigning equal weights to all stocks predicted to perform better than the median, finding that this portfolio provides higher returns and lower volatility than the index. Li et al. (2024) augment LSTM with Symbolic Genetic Programming (SGP) to create new features out of raw technical and fundamental indicators, predicting stock returns in the Chinese stock markets in Shenzhen and Shanghai. They find that this augmented model outperformed indices such as the CSI 300 and the CSI 500. Machine Learning and Reinforcement Learning for Portfolio Management: A variety of techniques to bolster traditional portfolio optimisation methods have been previously applied. For instance, Lu (2025) implements Ridge regression (Hoerl and Kennard, 1970) and Random Forests (Breiman, 2001) on historical prices and technical indicators for equities listed in the USA to predict future returns, subsequently applying mean-variance optimisation to the predicted returns to obtain the optimal portfolio. The constructed portfolios for both Ridge and Random Forests outperformed benchmarks. In a similar vein, Pinelis and Ruppert (2022) applied Elastic Nets and Random Forests to macroeconomic and financial data from 1927 to 2019, showing that these models provide higher returns without increasing volatility. Another technique from Kedia et al. (2018) involves using k-means clustering
(MacQueen, 1967) on the fundamental indicators of stocks comprising the BSE100 index of the Bombay Stock Exchange, to identify representative equities for clusters of companies with similar financial characteristics. They divide portfolio weights equally among the selected equities, and find that their portfolio outperformed the BSE100 index. Gu et al. (2025) develop an RL with Time Awareness, Short Selling and Attention mechanisms, trained on prices and technical data for stocks from the Dow Jones Industrial Average (DJIA). They report returns exceeding past neural networks based approaches. Jiang et al. (2024) implement a Deep Reinforcement Learning set-up using the Twin Delayed Deep Deterministic (TD3) agent (Fujimoto et al., 2018), aiming to optimize portfolio returns, on historical prices of stocks from the DJIA and the S&P 100 indices. They show that their set-up provides higher returns than traditional methods like mean-variance optimisation. Large Language Models (LLM) based: Zhang et al. (2024) develop a trading agent referred to as ‘FinAgent’ using the existing GPT-4 LLM from OpenAI. They provide prices, charts, company news, and financial reports, along with technical strategies and expert opinions to the LLM, to develop a three-part memory module consisting of “market intelligence” (a summary of all relevant data), lowlevel reflections on price movements, and high-level reflections on past trading decisions. They then apply this agent to five large-cap US equities and Ethereum, delivering higher returns than prior rulebased, ML, RL and LLM-based models. Research Gap: The literature suggests that there are extensive studies into the use of machine learning techniques in the context of financial markets in developed nations, such as the NASDAQ in the USA, while there is a dearth of comprehensive research into these techniques in India and other developing nations. Research based in developing nations would be particularly valuable given the unique regulatory environments, fast-paced growth, and young and expanding retail investor base. Moreover, applications using multiple modalities and types of data simultaneously, such as historical prices, technical indicators, financial fundamentals and social media and news texts and sentiments, are limited in the context of these markets. In this project, we aim to address the methodological and geographic gap by comparing the performance of multiple approaches to portfolio management, including Ridge regression, LLMs, LSTM and RL, utilising historical, fundamental, technical and sentiment-based features.
III. Methodology In this project, we focus on 50 traded equities on the National Stock Exchange, India which represent the NSE and constitute major indices such as the Nifty 50 (full list in appendix A). We utilise multimodal data for each of these equities for the year from June 2024 to June 2025, from the following data sources: 1. Historical Prices: The daily opening, high, low and closing prices, along with the volume traded (OHLCV for short) for each of the equities are obtained using the nsepython user-contributed open-source package (NSEPython Contributors, 2023). 2. Financial Indicators: Interim and annual financial statements, including income statements, cash flow statements and balance sheets, along with a range of metrics such as the Debt to Equity Ratio, Earnings per Share, Net Profit Margin, Revenue per Employee, and more, for each of the companies are obtained using the IndianAPI Stock API (IndianAPI, 2024). 3. News: Daily news headlines are obtained from the publication ‘The Economics Times’. News headlines and descriptions from a range of publications are obtained from the MarketAux API (MarketAux, 2024). 4. Social Media: Reddit posts and comments from selected communities that mention stock symbols of the companies are obtained from the official Reddit API (Reddit Inc., 2024). The selected communities are oriented around relevant discussions of financial news and trading strategies among members (full list in appendix B). X (formerly Twitter) posts from the official accounts of each company are obtained from twitterapi.io (TwitterAPI.io, 2024). We use the collected data to construct the following variables: 1. OHLCV data is used to construct selected autoregressive and technical indicators such as: a. Lagged closing price b. Lagged change in closing price c. Overnight gap (OpenPricet - ClosePricet-1) d. RSI 14 (Wilder, 1978): 100 - (100/(1 + RS)), where RS is the ratio of the average of the positive changes in closing price to the average of the negative changes in closing price over the past 14 days. e. Moving average of the closing price
f. Z-score of the closing price/volume: (mean of closing price or volume)/(standard deviation of closing price or volume) 2. Sentiment Indicators: News data is first matched to securities using fuzzy matching of the stock symbols with the headlines. This allows for the companies to be matched with the relevant headlines, while allowing for slight discrepancies based on naming conventions between the symbols and the headline texts. This is done using the fuzzywuzzy package in Python (SeatGeek, 2011). Next, we use FinBERT (Araci, 2019) to extract sentiment scores from each of the news items, reddit and twitter posts for each company. FinBERT (Financial Bidirectional Encoder Representations from Transformers) is a Natural Language Processing (NLP) model based on BERT (Devlin et al., 2019) for sentiment analysis specifically in the finance domain. This led to a total of 289 features for the 50 equities: 63 in OHLCV and derived features, 89 in technical indicators, 121 in financial indicators and 16 in sentiment indicators. Data is split into the first six months (June 2024 to January 2025) for training, and the next six months (January 2025 to June 2025) for testing. Having created the aforementioned features, we then consolidate the date by stock and align them to the trading days. After this, we move on to the architecture of our models. We first use some basic models, built on traditional portfolio management methods or simplistic strategies. These are as follows: 1. Equal weights: All 50 stocks are assigned equal weights and long positions throughout the time period. This is the simplest model, equivalent to buy-and-hold as a strategy. 2. Volatility adjusted equal weights: Portfolio weights as on date t are determined by the inverse of the volatility of each stock for the available observations up to t. 3. Past maximum return: The model optimises the weights assigned to each stock on date t to maximise historical return for the past 60 days. 4. Moving average crossover strategy: The model chooses to buy a long position in a stock on date t if the moving average of its closing price over the past 10 days is more than the moving average over 20 days. Portfolio weights are allocated equally between the stocks that are chosen to be bought. 5. Momentum-based portfolio: Portfolio weights as on date t are determined by the cumulative return of each stock over the past 20 days.
6. Black-Litterman model (Black and Litterman, 1992): First, the covariance matrix for the returns of the stocks as on date t is calculated using Ledoit-Wolf shrinkage (Ledoit and Wolf, 2004) using data from the past 60 days. The market strategy is assumed to be the equal weights strategy, while the investor views are assumed to be the momentum-based strategy, as defined above. The expected returns for each stock are then calculated as: r = [τΣ-1 + PT Ω-1 P]-1 × [τΣ-1π + PT Ω-1 Q], where τ is a scaling factor, Σ is the estimated covariance matrix, P is the views matrix (i.e, which stock is affected by which view, in our simplified model this is the identity matrix, which can be interpreted as each stock’s momentum is viewed to only affect itself), Ω is the matrix of the confidence in the investor views (for simplicity, this is taken to be the diagonal matrix of the individual stocks’ variance), Q is the vector of the investor views (momentum of each stock), and π is the market equilibrium returns (Σ multiplied with the equal weights vector). The weights are finally chosen to maximise the objective function: wT r - (λ/2) × wT Σ w, where w is the vector of the assigned weights, with the constraints that the weights are all positive and sum to 1. 7. Minimum variance portfolio: The correlation matrix Σ as on date t is calculated using data from the past 60 days. Portfolio weights are chosen to minimize the objective function: wT Σ w, where w is the vector of the assigned weights, subject to the constraints that the weights are all positive and sum to 1. This strategy only aims to minimise risk, without considering returns. 8. Technical analysis strategy: This strategy calculates three technical indicators using the closing price of each stock: Relative Strength Index (RSI 14) (Wilder, 1978), Bollinger Bands position (Bollinger, 1992) and Moving Average Convergence Divergence (MACD) (Appel, 1979). RSI is calculated as the final value of the vector obtained using the following formula: 100 - (100/(1 + RS)), where RS is a vector of the ratio of positive changes in closing price to the negative changes in closing prices over the previous 14 days. The Bollinger Bands are defined as the mean ± 2 standard deviations of the closing prices over the previous 20 days. Bollinger Bands position is then calculated using the following formula:
BB_position = (Current Closing Price - BB_lower) / (BB_upper - BB_lower) MACD is calculated as the difference between the exponentially weighted means of the closing prices of the previous 12 and 26 days. Finally, the signals from the RSI, Bollinger Bands position and MACD are averaged to form one signal, which is used to determine the portfolio weights. 9. Sentiment based strategy: This strategy takes the average of the news, Reddit and Twitter sentiment scores generated using FinBERT, as discussed above. The stocks with positive sentiment scores are selected, and portfolio weights are assigned proportional to the scores. These are the models we study: Ridge regression + Rule-based model: First, we use a linear model, with Ridge regularisation (Hoerl and Kennard, 1970), to predict the logarithm of the change of each individual equity’s closing price from day t - 1 to t, using all the features described above for the period from t - 120 to t - 1. Next, we estimate the covariance matrix of the closing price using data from t - 90 to t - 1 using the Ledoit-Wolf shrinkage estimator (Ledoit and Wolf, 2004). The top 15 stocks with the highest predicted returns are selected. Weights are chosen to maximise returns on t, penalised by the risk (estimated as the product of the weights vector with the estimated covariance matrix), turnover and transaction cost. This optimization problem is solved numerically using the OSQP solver (Stellato et al., 2020). The mathematical representation of the optimization problem is: Maximise μT w - λR (wT Σ w) - λT ||Δw|| - (c |Δw|) with respect to w, Where μ is the predicted next day log returns, w is the vector of the assigned weights, λR is a riskaversion penalty (set to 0.5%), Σ is the estimated covariance matrix, λT is a turnover penalty (set to 0.1%), ||Δw|| is the sum of the absolute values of the changes in weights from t - 1 to t, c is the transaction cost (set to 0.5%). The constraints are that the weights sum to 1, and each of the weights assigned to a stock lies between 0 and 20%, to avoid excess allocation to one individual stock.
Figure 3 reports the risk metrics for the Black-Litterman model. The Sharpe ratio is highly variable, ranging from -3 to 5. Volatility shows a general downward trend, going from almost 26% at the start of the trading period to about 15% at the end. The drawdown shows a few major dips in portfolio value around March and the end of May. These metrics suggest that the portfolio constructed by the BlackLitterman model is highly volatile.
Ridge Regression + Rule-based model Fig. 4: Summary of Ridge Regression + Rule-based model Figure 4 reports the performance summary for the Ridge regression + Rule-based model. The portfolio value rises from January to June and finishes at a terminal value of about 1.192 million on an initial investment of 1 million, which corresponds to a total return near 19.2%. Risk is moderate with volatility near 13.5% and a maximum drawdown of about 6.9 percent. The histogram depicts a mean daily return of about 0.12% and a distribution portraying frequent small gains. Turnover averages about 0.6% per day, which is consistent with small daily adjustments to the portfolio. These panels together indicate that the model generated steady compounding with shallow and recoverable drawdowns.
Fig. 5: Accuracy of Ridge Regression The confusion matrix in Figure 5 summarises the quality of the model’s directional predictions at the daily horizon, i.e, whether the closing price would go up or down on the next day. Accuracy is 0.820, precision is 0.871, recall is 0.871, and the F1 score is 0.871. These values suggest that the signal is both selective and consistent. The small difference between the error types implies that misclassifications are not skewed toward only one side of the market.
Fig. 6: Ridge Regression + Rule-based model risk metrics Figure 6 shows rolling risk diagnostics. The 30-day Sharpe ratio climbs above one by late March, peaks near eight in mid-April, and then decays toward one by early June. The daily drawdown plot highlights a single deep valley in mid-March that is fully recovered by April. Volatility ranges between 9 and 18 percent, highest during mid-May. Together these diagnostics indicate that the model scaled risk appropriately over time and avoided persistent periods of negative returns.
Fig. 7: Average holdings distribution for Ridge Regression + Rule-based model Figure 7 suggests that this model concentrates capital in a small set of liquid NIFTY names while maintaining broad sector coverage. The model chooses equities from sectors ranging from energy (NTPC, Reliance), to fast-moving consumer goods (Tata Consumer, Britannia), to banking (IndusInd Bank, Axis Bank). The top five positions (NTPC, HDFC Life, Shree Cement, Tata Consumer and IndusInd Bank) account for about 64.5% of average portfolio weight, indicating stable allocations in top performing companies. This also suggests that the model chooses to avoid excessive exposure in similar stocks, prefers differentiation, and avoids high volatility. LLM-based Model
Fig. 8: Summary of LLM Model Figure 8 shows a steady equity climb from January to early June with only brief pullbacks for the LLM-based model. The return reaches roughly 11.9% by the end of the window, depicting smooth compounding. The daily-return histogram is narrow and centred slightly above zero, indicating small but frequent positive days and very few large losses. Drawdowns are shallow and short lived, mostly contained within 0.4%, which matches the visually small dips in the top panel. Together these patterns point to a conservative execution style with low volatility and low turnover, and compounds through a high daily hit-rate of small gains.
Fig. 9: LLM-based model risk metrics Figure 9 shows the risk metrics of the LLM-based model, which points to the low volatility depicted in Figure 8. The Sharpe ratio remains consistently high as a result of the steady increases in the portfolio value without large deviations. The drawdown is also marginal after initial fluctuations in January. Fig. 10: Average holdings distributions for LLM-based model
Figure 10 shows that weights are well spread across the top ten positions, with sectors spanning electronics and defence (BEL), life insurance (SBI Life), healthcare services (Max Healthcare), retail (Trent), financials (Jio Finance, Shriram Finance), energy and infrastructure (ONGC, Adani Enterprises), and steel (JSW Steel). The dispersion of weights signals that the agent is not overconcentrated. LSTM-based model We discuss the results for the three portfolio allocation methods under the LSTM based model separately. First, we present the Top-K Long positions model, second the Risk Adjusted Weights model, and third the Equal Weights model. Fig. 11: Summary of LSTM + Top-K Long positions The long-only LSTM on the top-K stocks (as shown in Figure 11) compounds steadily across the test window. Total return is 9.39% and a Sharpe ratio of 1.22. The portfolio value shows an early drawdown into February and March, followed by a persistent recovery that carries through to early June. The drawdown panel peaks near 9.1% and then recedes, which indicates losses were contained and recoverable. The daily-returns histogram centers slightly above zero with a sample mean of 0.09%
and thin right-tail outliers. Turnover is low, generally below 4% per day, which is consistent with a patient profile rather than rapid rebalancing. This is explained by the allocation only changing if the prediction of the Top 10 stock changes, and remaining the same otherwise. This activity level is friendly to realistic costs and makes the strategy straightforward to execute. Fig. 12: LSTM + Top-K Long positions risk metrics Figure 12 shows the risk metrics for the LSTM top-10 long only strategy, which align with the performance summary in Figure 11. As the model recovers from the dip in February and March, the Sharpe ratio increases past 1. Rolling volatility shows an increasing trend, however it remains commensurate with the returns from the model.
Fig. 13: Summary of LSTM + Risk-Adjusted Weights Model The risk-adjusted LSTM in Figure 13 finishes the window with a total return of 18.8% and a Sharpe ratio of 2.77. The portfolio value shows a quicker recovery from early losses and a more stable climb into June, compared to the top-K long positions strategy. The drawdown profile peaks near 6.9%, which is lower than the other LSTM strategies. The daily-return histogram is centered slightly above zero (mean 0.11%) with a visibly thinner left tail, consistent with the reduction in maximum drawdown. Daily turnover remains modest, below 3.5%, which keeps transaction costs low. As seen in Figure 14, the volatility remains lower than 12%, and the Sharpe ratio rises from April onwards. These patterns indicate that risk adjusted portfolio weight allocations curbed downside while preserving upside.
Fig. 21: PPO + LSTM model risk metrics Figure 21 reflects the higher risk associated with the PPO + LSTM model. The Sharpe ratio falls but then recovers to above 1 during March and May, corresponding to the recovery from large drawdowns exhibited in Figure 20. Rolling volatility remains above 18% for the majority of the testing period, suggesting a high risk, high reward approach. Comparison of Models We now discuss the comparative results of all the models we have tested in this project. Figure 22 and Table 1 depict the performance of all 16 strategies implemented, with 9 baseline strategies and 7 machine learning based strategies. We focus on Figure 21, which depicts the four main strategies that are the focus of this paper: Ridge regression, LLM-based, LSTM, and PPO, with Black-Litterman for reference. The PPO model with an LSTM policy function had the highest return at 19.38%, followed by Ridge at 19.17%, LSTM with Risk-Adjusted weights at 18.79%, and LLM-based at 11.93%. However, we note that while other strategies have large variations in portfolio value over time, the LLM-based model steadily rises, causing it to have the highest Sharpe ratio at 3.99, followed by LSTM with Risk-Adjusted weights at 2.77, Ridge at 2.23 and PPO with LSTM at 1.89.
In general, our findings indicate that the models have different strengths. While PPO provides high returns, it also has very high volatility. This is expected as the reward function is specified to incentivise higher returns, with no controls for risk. Ridge and LSTM provide comparable returns at lower volatility. This is by virtue of their portfolio optimisation mechanism, with the Ridge model assigning weights to maximise returns and minimise risk, while the Risk-Adjusted weights for the LSTM model assigns weights inversely proportional to volatility. The LLM based model provides lower returns, but very low volatility, with the LLM choosing minimal transactions after the original portfolio allocation. It is important to note that the PPO model is extremely intensive to train, in terms of computational power, memory and time. The LSTM model is moderately intensive, while the Ridge model, being a linear regularised regression, is least demanding. Fig. 22: Top Performing Strategies Across Model Categories
Fig. 22: Performance comparison between all strategies
Table 1: Summary of Performance of All Models Rank Strategy Total Return (%) Annualized Return (%) Annualized Volatility (%) Sharpe Ratio 1 PPO-LSTM 19.38 52.99 28.01 1.89 2 Ridge Regression 19.17 52.33 23.44 2.23 3 LSTM - Risk Adjusted 18.79 51.16 18.49 2.77 4 LSTM - Equal Weight Positive 15.67 41.82 20.59 2.03 5 LLM-based 11.93 31.06 7.79 3.99 6 Black-Litterman 11.78 30.64 27.83 1.1 7 Equal Weight 10.52 27.15 16.42 1.65 8 LSTM - Top 10 Long 9.39 24.05 19.69 1.22 9 PPO-MLP 5.41 13.49 20.84 0.65 10 Sentiment-based 4.34 10.73 4.97 2.16 11 Volatilityadjusted Equal Weight 3.94 9.72 13.55 0.72 12 Momentum 3.67 9.05 17.57 0.52 13 Minimum Variance 0.01 0.02 0.46 0.05 14 Maximum Return 60day -8.71 -19.64 17.19 -1.14 15 Moving Average Crossover -16.26 -34.69 13.39 -2.59 16 Technical Analysis -52.61 -83.34 8.23 -10.13
V. Conclusion This project provides a comparison of four machine learning methods in portfolio management in the Indian equities market, while using data from historical prices, financial reports, technical indicators, news headlines and social media posts. The findings show gains over traditional portfolio optimisation techniques such as the Black-Litterman model in all methods. However, we find that RL-based models, in spite of their high computational requirements, do not provide significant gains over a simple regularised linear regression like Ridge. This finding may be of interest to researchers and practitioners looking to apply machine learning methods, especially for higher frequency and faster training and implementation. Our study has some limitations. The shorter time period being studied may raise concerns about the applicability of our findings to other market regimes and cycles. In this regard, studying a wider time period with a larger range of securities may be valuable. The black-box nature of neural networks and reinforcement learning models limits the explainability of our strategies. Further work into the importance of specific features and the influences of market changes on trading decisions would be valuable. Finally, taking into account more constraints on liquidity, transaction costs and tax liabilities could improve the realism of our models, and also move towards real world productionisation for the best performers.
VI. References Appel, G., 1979. “The Moving Average Convergence-Divergence Trading Method.” Scientific Investment Systems. Araci, D., 2019. "FinBERT: Financial Sentiment Analysis with Pre-trained Language Models," arXiv preprint arXiv:1908.10063. Black, F. and R. Litterman, 1992. "Global Portfolio Optimization," Financial Analysts Journal 48(5), pp. 28-43. Bollinger, J., 1992. "Using Bollinger Bands," Stocks & Commodities 10(2), pp. 47-51. Breiman, L., 2001. "Random Forests," Machine Learning 45(1), pp. 5-32. Business India, 2024. "The Rise and Rise of Retail Investors," October 26, 2024. Available at: https://businessindia.co/magazine/cover-feature/the-rise-and-rise-of-retail-investors Chaudhary, R., 2025. "Sentiment-Aware LSTM Framework for Technology Stock Prediction," arXiv preprint arXiv:2505.05325v1. Devlin, J., M. Chang, K. Lee, and K. Toutanova, 2019. "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding," Proceedings of NAACL-HLT 2019, pp. 4171-4186. Fjellström, C., 2022. "Long Short-Term Memory Neural Network for Financial Time Series," arXiv preprint arXiv:2201.08218. Fujimoto, S., H. van Hoof, and D. Meger, 2018. "Addressing Function Approximation Error in ActorCritic Methods," Proceedings of the 35th International Conference on Machine Learning, PMLR 80, pp. 1587-1596. Hochreiter, S. and J. Schmidhuber, 1997. "Long Short-Term Memory," Neural Computation 9(8), pp. 1735-1780.
Hoerl, A.E. and R.W. Kennard, 1970. "Ridge Regression: Biased Estimation for Nonorthogonal Problems," Technometrics 12(1), pp. 55-67. Hutto, C.J. and E. Gilbert, 2014. "VADER: A Parsimonious Rule-based Model for Sentiment Analysis of Social Media Text," Proceedings of the Eighth International AAAI Conference on Weblogs and Social Media, pp. 216-225. IBEF (India Brand Equity Foundation), 2024. "Rise of Retail Investors and Domestic Funds Driving Growth in India," September 4, 2024. Available at: https://ibef.org/blogs/rise-of-retail-investors-anddomestic-funds-in-india IBEF (India Brand Equity Foundation), 2025. "Investment in India: Market Trends, Sectors & Opportunities." Available at: https://ibef.org/economy/investments ICICI Direct, 2025. "India's 2024 Share Market Summarized: Know All Details," January 3, 2025. Available at: https://www.icicidirect.com/research/equity/finace/summary-of-indian-share-market-2024 IndianAPI, 2024. Indian Stock Market API Documentation. Available at: https://indianapi.in/ Jegadeesh, N. and S. Titman, 1993. "Returns to Buying Winners and Selling Losers: Implications for Stock Market Efficiency," Journal of Finance 48(1), pp. 65-91. Jiang, Z., D. Xu, and J. Liang, 2024. "A Deep Reinforcement Learning Framework for the Financial Portfolio Management Problem," Economic Modelling 135, pp. 106887. Kedia, S., K. Mehra, and S. Sharma, 2018. "Stock Market Portfolio Selection using K-Means Clustering Algorithm," 2018 International Conference on Advances in Computing, Communications and Informatics (ICACCI), pp. 1333-1338. Ledoit, O. and M. Wolf, 2004. "Honey, I Shrunk the Sample Covariance Matrix," Journal of Portfolio Management 30(4), pp. 110-119. Li, Y., W. Zheng, and Z. Zheng, 2024. "Deep Learning for Stock Prediction Using Numerical and Textual Information," Scientific Reports 14, Article 1234.
Lu, X., 2025. "Machine Learning Enhanced Mean-Variance Framework for Portfolio Optimization," Proceedings of ICDEBA 2024. Atlantis Press. MacQueen, J., 1967. "Some Methods for Classification and Analysis of Multivariate Observations," Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability 1, pp. 281297. Markowitz, H., 1952. "Portfolio Selection," Journal of Finance 7(1), pp. 77-91. MarketAux, 2024. MarketAux News API Documentation. Available at: https://marketaux.com/ Meher, B.K. and R. Mishra, 2024. "Risk-Adjusted Portfolio Performance Analysis using BlackLitterman Model: Evidence from Indian Equity Markets," Australasian Accounting, Business and Finance Journal 18(2), pp. 45-67. Metropolis, N. and S. Ulam, 1949. "The Monte Carlo Method," Journal of the American Statistical Association 44(247), pp. 335-341. Mistral AI, 2023. Mistral Medium: Technical Documentation. Available at: https://mistral.ai/ News4Masses, 2025. "India's NSE Crosses 22 Crore (220 Million) Total Investor Accounts," April 14, 2025. Available at: https://news4masses.com/nse-investor-growth-2025/ NSEPython Contributors, 2023. NSEPython: Python Library for NSE India. Available at: https://github.com/jugaad-py/jugaad-data OpenAI, 2023. "GPT-4 Technical Report," arXiv preprint arXiv:2303.08774. Outlook Business, 2025. "The Retail Investor Revolution: How 19 Crore Demat Accounts Reshape India's Market," July 16, 2025. Available at: https://www.outlookbusiness.com/markets/the-retailinvestor-revolution-how-19-crore-demat-accounts-reshape-indias-market Pelluri, K., 2025. "Momentum-Based Investment Strategies in Indian Mid-Cap Equity Markets," SSRN Working Paper 5116091.
Raffin, A., A. Hill, A. Gleave, A. Kanervisto, M. Ernestus, and N. Dormann, 2021. "Stable-Baselines3: Reliable Reinforcement Learning Implementations," Journal of Machine Learning Research 22(268), pp. 1-8. Reddit Inc., 2024. Reddit API Documentation. Available at: https://www.reddit.com/dev/api/ Schulman, J., P. Moritz, S. Levine, M. Jordan, and P. Abbeel, 2015. "High-Dimensional Continuous Control Using Generalized Advantage Estimation," arXiv preprint arXiv:1506.02438. Schulman, J., F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, 2017. "Proximal Policy Optimization Algorithms," arXiv preprint arXiv:1707.06347. SeatGeek, 2011. FuzzyWuzzy: Fuzzy String Matching in Python. Available at: https://github.com/seatgeek/fuzzywuzzy Sharpe, W.F., 1966. "Mutual Fund Performance," Journal of Business 39(1), pp. 119-138. Stellato, B., G. Banjac, P. Goulart, A. Bemporad, and S. Boyd, 2020. "OSQP: An Operator Splitting Solver for Quadratic Programs," Mathematical Programming Computation 12(4), pp. 637-672. TwitterAPI.io, 2024. Twitter API Service Documentation. Available at: https://twitterapi.io/ Wilder, J.W., 1978. New Concepts in Technical Trading Systems. Trend Research. Zhang, H., X. Liu, and T. Wang, 2024. "FinAgent: A Multimodal Foundation Agent for Financial Trading," arXiv preprint arXiv:2402.18485. Appendix A: List of Stocks Studied RELIANCE TCS HDFCBANK INFY ICICIBANK HINDUNILVR BEL SBIN BHARTIARTL KOTAKBANK
LT ASIANPAINT AXISBANK MARUTI ULTRACEMCO TITAN SHRIRAMFIN NESTLEIND ONGC ETERNAL NTPC ITC BAJAJFINSV SUNPHARMA JIOFIN WIPRO HCLTECH JSWSTEEL TATAMOTORS ADANIPORTS BRITANNIA SHREECEM TATACONSUM BPCL INDUSINDBK GRASIM ADANIENT DRREDDY COALINDIA HEROMOTOCO IOC UPL CIPLA TATASTEEL HDFCLIFE SHREECEM BAJAJ-AUTO SBILIFE MAXHEALTH TRENT Appendix B: List of Reddit Communities Referred r/IndianStockMarket r/IndiaInvestments