Financial-Firm 10-K Access Claims
SEC disclosure text pipeline linked to benchmark-adjusted return windows.
Research pipeline for testing whether financial and fintech-related firms that use 10-K language about democratizing finance, financial inclusion, lower barriers, retail investors, underserved borrowers, or institutional-quality products for individuals show different subsequent stock performance.
The strength of the project is the audit trail. It does not treat raw phrase hits as evidence. It builds filing-level treatments from 10-K sections, validates conservative text measures, links securities to CRSP-style returns, computes forward return windows, and documents why the current evidence should be read as pilot evidence rather than a causal or publication-ready result.
Pilot empirical-finance research archive; useful as an auditable text-measure and return-pipeline prototype
What It Does
- Builds a financial and fintech firm universe, filing index, SEC filing download workflow, and section extraction pipeline for Item 1, Item 1A, and Item 7 text.
- Constructs and audits phrase-based treatment measures for democratization, inclusion, lower-barrier, retail-availability, and underserved-borrower language.
- Rejects broad or weak classifiers when validation evidence does not support them, preserving a more conservative filing-level treatment.
- Imports and validates WRDS/CRSP security links, daily returns, market benchmarks, and 1-, 3-, and 5-year compounded return windows.
- Separates raw returns from winsorized and benchmark-adjusted outcomes so later analysis remains auditable.
Outputs
- Analysis tables for baseline estimates, inference diagnostics, winsorization thresholds, and sample-support diagnostics.
- SQL files for WRDS/CRSP linking, daily-return pulls, market benchmarks, and schema setup.
- Checkpoint logs and quality reports documenting ingestion, section extraction, phrase hits, classification audits, treatment construction, return windows, and baseline modeling.
- Methodology notes covering research design, pre-analysis planning, return methodology, outlier policy, and model specification.
Limitations
The repository is explicit that the evidence is observational and not causal. The main limitation is sample construction: the pilot firm universe was built from a current-listed SEC ticker feed rather than a point-in-time CRSP/Compustat/security-master universe. That creates survivorship risk and limits how strongly the return results should be interpreted.
The current baseline result is best read as a careful pilot pipeline rather than a strong null or return-predictability claim. The 1-year result is the most interpretable, while the 3-year and 5-year windows are too imprecise for strong conclusions.