The field of systematic equity investing generates an enormous amount of research. Papers documenting new predictors of stock returns appear continuously. Factor libraries expand to include dozens and then hundreds of candidate signals. Machine learning and language models open new routes to extracting patterns from data that earlier methods could not reach. The pace of production is not matched by a corresponding clarity about what it all means. A researcher facing a historical return pattern that looks attractive still has to answer the same hard questions: Is this real? How confident should I be? What would happen to the edge if more capital pursued it? How much of the apparent performance would survive live trading? When would I retire the signal and why?
This book is about how to answer those questions. It is not a catalog of known factors. It does not rank strategies by their historical performance or explain which sectors offer the best opportunities at any given moment. Its subject is the research process itself: the sequence of decisions a quantitative researcher makes when taking a candidate signal from its first formulation through a disciplined validation and into live use, and the specific ways each of those decisions can go wrong.
The book is structured around fourteen questions. Each chapter addresses one of them, because each question names a genuine decision point that every systematic signal must pass through. What are you predicting, and have you stated it precisely enough to be tested? Can you trust the data underlying that prediction? How much evidence genuinely survives the model-selection process you used to find the signal? Where does the signal’s edge actually live in the return distribution? What will it cost to trade? What unintended risks are embedded in it? When should it be retired? These are not rhetorical or philosophical questions. They are operational ones with specific, checkable answers that determine whether a candidate signal is worth deploying and how it should be managed once it is.
This book organizes systematic equity research around the decisions that produce a result: defining its target, establishing its information set, testing it, and deciding whether it merits capital. The same signal, evaluated with more or less rigor at each stage, can produce results that appear entirely different without any change in the underlying data. A factor with a plausible economic story and an attractive backtest can be a genuine alpha source or an artifact of data mining, survivorship, look-ahead bias, or estimation noise. Which of these it is cannot be determined by reading the final performance number. Assessing those explanations requires examining the decisions that produced the number; even a careful audit leaves uncertainty about future performance.
The presentation here is organized around reasoning, not around results. Each chapter builds a framework for making a specific class of decision correctly and then examines the ways that class of decision goes wrong when the framework is not applied. The examples throughout are drawn from academic and practitioner research, including papers, books, and research reports, chosen because they illustrate a research principle clearly rather than to survey the literature comprehensively. The goal is to give the reader a durable way of thinking about each problem, not a catalog of findings to memorize.
The book is aimed primarily at graduate students in quantitative finance, financial engineering, statistics, or economics who are learning systematic investment research for the first time, and at quantitative practitioners who want a more systematic account of the research methodology underlying what they do. Data scientists entering investment management from outside finance will find it useful as a structured introduction to the specific challenges of financial data and financial research that are not present in other domains. The technical level assumes comfort with statistics and hypothesis testing, basic finance, and data analysis in a programming environment. Familiarity with machine learning methods will be helpful for several chapters, particularly the discussions of model selection and alternative data. Advanced mathematics, derivatives pricing, stochastic calculus, and portfolio optimization theory are not required.
A second reader profile worth naming is the experienced systematic practitioner who developed research skills primarily through practice. Such readers often have strong empirical instincts about specific problems they work on daily but have not had occasion to work through the methodological foundations in a coherent sequence. The book is designed to be useful to them as a structured framework for the decisions they already make, filling in the conceptual grounding that practice alone rarely provides. These readers will find the chapters on backtest design, signal combination, and governance most directly applicable to their current work, but chapters on data lineage and prediction target specification sometimes contain the explanations for problems they have observed empirically without a clear account of why those problems occur.
How the Book Is Organized
The fourteen chapters fall naturally into three phases, and reading them in order is strongly recommended.
The first five chapters establish the research infrastructure. Chapter 1 treats the prediction target as a design object: before any feature is constructed, the researcher must specify what outcome is being predicted, for which entity population, over which horizon, relative to which benchmark, and at which forecast origin. These choices determine what data is needed, what models can be estimated, and what a positive result would mean. Chapter 2 develops the temporal discipline that separates valid from invalid research design: the requirement that mechanisms, evaluation methods, and execution assumptions be committed to before the confirmatory evaluation is observed, and the five distinct information clocks that each input to a model passes through on its way from the real world to a decision. Chapter 3 addresses data trustworthiness: point-in-time data, entity crosswalks, coverage funnels, and the distinction between when data was recorded and when it was available. Chapter 4 turns to feature construction, asking whether a particular transformation of a data source actually measures the economic object it claims to measure, and what mathematical properties of a feature determine whether it contains useful predictive content. Chapter 5 addresses evidence quality: the four tiers of evidence (discovery, validation, final test, and live), the statistical traps that arise when models are selected over many candidates, and the architecture of a walk-forward evaluation that limits the number of free choices made after observing results.
The next four chapters develop signal characterization. Chapter 6 examines what the rank information coefficient misses about a cross-sectional signal: payoff shapes, long-side and short-side decompositions, conditional structures, and how each diagnostic changes the appropriate portfolio architecture. Chapter 7 covers event-driven research: how to design the event clock, manage the three windows around each event (pre-announcement, announcement, post-announcement), and handle the statistical complications of event timing endogeneity and event clustering. Chapter 8 addresses text, network, and alternative data, developing a validity framework built on what a data source measures rather than how sophisticated its production pipeline is, and examining the specific contamination risks that language model research inherits from the classical problems the field already knew about. Chapter 9 discusses regime-conditional signals: when signals work better in some market states than others, how to design conditioning frameworks that avoid hindsight at the classification layer, and how to govern a research program that makes claims conditional on states that cannot be perfectly identified in real time.
The final five chapters cover portfolio integration. Chapter 10 distinguishes statistical evidence from economic evidence, showing why a statistically significant historical information coefficient is not the same as a tradeable investment opportunity once breadth, constraint, and costs are accounted for. Chapter 11 covers turnover, transaction costs, borrow expense, and capacity, treating them not as haircuts applied after the fact but as structural properties of a strategy’s trading behavior that must be modeled inside the backtest. Chapter 12 addresses risk contamination: the tendency of raw signals to carry embedded factor exposures that were not part of the intended forecast, and the toolkit for removing those exposures while seeking to preserve as much genuine predictive content as possible. Chapter 13 covers signal combination: how to assess whether a new component contributes incremental information to an existing composite, how to calibrate components with different scaling conventions, and how to freeze the combination framework so that live revisions require the same evidentiary standard as the original deployment. Chapter 14 closes with research governance: monitoring a live signal’s returns alongside costs, exposures, and data quality; distinguishing the three sources of live signal underperformance; choosing among continued monitoring, temporary suspension, repair, substitution, and retirement according to the diagnosis and expected net value; and building the organizational structures and documentation practices that allow a governance program to function rather than merely exist.
The chapters are designed to be read in order because each one builds on decisions made in the previous ones. Chapters 1 through 5 are the foundation, and the later chapters assume their vocabulary and disciplines. A practitioner who wants a shorter engagement with the material will find Chapters 1, 2, 5, 10, 11, 13, and 14 the most immediately applicable to the questions that arise in daily research work, but should be aware that these chapters draw on frameworks built in the chapters they skip.
This book does not tell the reader which factors to invest in. It does not evaluate the current return prospects of value, momentum, profitability, or any other known factor family, and it does not recommend strategies. A factor with a long academic track record may be used here and there as an example of a research discipline point, not as an endorsement. The book does not cover portfolio construction beyond the signal-level decisions described in Chapters 11, 12, and 13. Optimization, liability-aware investing, transaction cost minimization at the portfolio level, and multi-asset allocation are substantial fields in their own right. The book covers the portfolio-construction decisions needed to evaluate and implement signals; it is not a comprehensive treatment of optimization, liability management, or multi-asset allocation.
A Note on Sources
This book is based on published research, including papers and books. It does not reveal proprietary strategies, trading signals, or firm-specific methodologies. Where sources describe experiments using proprietary datasets or institutional research, this book discusses only the methods and findings disclosed in those sources.
This book grew out of lecture notes developed for a graduate course in quantitative finance methods. That origin shapes its structure and its tone. Lecture notes are organized around understanding rather than comprehensive coverage; they teach the reasoning rather than exhaustively cataloging the literature. The emphasis throughout is on frameworks and ways of thinking, not on compiling results. Each chapter ends with a short Further Reading section for students who want to explore the underlying literature more deeply. Those suggestions point to academic papers and books that treat the chapter’s main subject at greater length; they are a starting point for further study.