Research & Methodology
The Science
Behind the Games
Behavioral economics. Reinforcement learning. DSGE modeling. Cognitive psychology. UI research. These are not decorative influences — they are architectural constraints.
01
Abstract
This document describes the methodological foundations underlying Screwcap Games, LLC's portfolio of browser-based games. Our design philosophy is predicated on the proposition that entertainment and education are not in tension — that genuinely calibrated challenge, grounded in real economic and behavioral models, produces both higher engagement and measurable transfer of skill.
The Screwcap stack draws from five disciplines: reinforcement learning (DoubleFives PPO agent), DSGE modeling (www.chairmanmode.app), behavioral economics (Gold Digger), cognitive psychology (difficulty scaling), and UI/UX research (interface friction). These are not marketing claims. They are architectural constraints embedded in how each game is built.
02
AI Methodology: DoubleFives PPO Agent
The AI powering DoubleFives is a Proximal Policy Optimization (PPO) agent trained via deep reinforcement learning in a four-player adversarial domino environment. PPO belongs to the same algorithm family as OpenAI Five and the predecessor approaches to AlphaGo — it is among the most battle-tested policy gradient methods in the field.
The agent uses an Actor-Critic architecture with a 3-layer MLP policy network. Seven distinct AI personalities were developed by training separate policy heads with modified reward functions — introducing different weightings on aggressive play, defensive blocking, and partner cooperation. These produce genuinely distinct behavioral profiles, not cosmetic variation.
Model weights are exported via ONNXand run entirely in-browser. No game moves are transmitted to a server. The AI runs locally on the player's device.
Primary Citations
- [→]Schulman et al. (2017). "Proximal Policy Optimization Algorithms." arXiv:1707.06347. — The foundational PPO paper. Our implementation follows the clipped surrogate objective.
- [→]OpenAI et al. (2019). "Dota 2 with Large Scale Deep Reinforcement Learning." arXiv:1912.06680. — OpenAI Five PPO scaling.
- [→]Vinyals et al. (2019). "Grandmaster level in StarCraft II using multi-agent reinforcement learning." Nature, 575, 350–354.
03
Economic Model Design: www.chairmanmode.app
www.chairmanmode.app is built on a Dynamic Stochastic General Equilibrium (DSGE) model — the workhorse framework of modern monetary policy analysis, used by the Federal Reserve, ECB, and major academic institutions. It implements multiple historically calibrated scenarios — 1929, 2008, and the Volcker Shock among them — with parameters drawn from FRED.
It implements the Taylor Rule as a policy benchmark:
Primary Citations
- [→]Smets, F. & Wouters, R. (2007). "Shocks and Frictions in US Business Cycles: A Bayesian DSGE Approach." American Economic Review, 97(3), 586–606.
- [→]Taylor, J.B. (1993). "Discretion versus policy rules in practice." Carnegie-Rochester Conference Series, 39, 195–214. — The Taylor Rule.
- [→]Galí, J. & Gertler, M. (1999). "Inflation dynamics: A structural econometric analysis." Journal of Monetary Economics, 44(2), 195–222.
- [→]Woodford, M. (2003). Interest and Prices: Foundations of a Theory of Monetary Policy. Princeton University Press. — The canonical New Keynesian DSGE treatment.
04
Behavioral Economics in Game Design
Gold Digger's prediction market framing exploits the well-documented gap between objective probability and subjective probability weighting identified by Kahneman and Tversky. Players systematically overweight small probabilities — the game makes this miscalibration visible through immediate feedback. The pedagogical goal is probability calibration.
DoubleFives' difficulty design is built the same way, one layer down: seven policy heads, each a separate PPO training run with a different reward weighting on aggressive scoring, defensive blocking, and partner cooperation, give seven genuinely distinct opponents rather than one AI throttled by a difficulty slider. That range is what lets the game sit each player at the boundary between anxiety and boredom that Csikszentmihalyi's Flow theory describes, across two, three, and four-seat play, rather than tuning a single opponent down until it stops being a real game.
Primary Citations
- [→]Kahneman, D. & Tversky, A. (1979). "Prospect Theory: An Analysis of Decision under Risk." Econometrica, 47(2), 263–292.
- [→]Kahneman, D. (2011). Thinking, Fast and Slow. Farrar, Straus and Giroux.
- [→]Thaler, R.H. & Sunstein, C.R. (2008). Nudge: Improving Decisions About Health, Wealth, and Happiness. Yale University Press.
- [→]Csikszentmihalyi, M. (1990). Flow: The Psychology of Optimal Experience. Harper & Row.
05
Rating Methodology
DoubleFives agents are rated with Glicko, not classic Elo. Elo assumes every rating is equally certain; Glicko carries a rating deviation alongside the rating itself, which matters here because a new policy head plays thousands of games against the established roster within days — treating those early results as exactly as reliable as a rating with a year of games behind it would be the wrong kind of confident.
As of the latest measured checkpoint, the strongest neural policy head rates meaningfully above the scripted baselines it trains against (greedy-scoring, tile-counting, and random-play bots) — the gap the game means when it says the AI “counted better.” The board is public and updates as training continues; see the live figures on the homepage Lab section.
06
UI/UX Design Principles
Screwcap interfaces minimize extraneous cognitive load while preserving germane cognitive load — the effortful engagement required to learn the underlying model. A player struggling with TheChair's rate decisions should be struggling with the macroeconomics, not the UI.
Primary Citations
- [→]Nielsen, J. (1994). "10 Usability Heuristics for User Interface Design." Nielsen Norman Group.
- [→]Sweller, J. (1988). "Cognitive Load During Problem Solving: Effects on Learning." Cognitive Science, 12(2), 257–285.
- [→]Fitts, P.M. (1954). "The information capacity of the human motor system." Journal of Experimental Psychology, 47(6), 381–391.
- [→]Miller, G.A. (1956). "The magical number seven, plus or minus two: Some limits on our capacity for processing information." Psychological Review, 63(2), 81–97.
07
Raw Training Metrics
Live from the same feed as the homepage Lab section — not a snapshot pinned to one training run.
08
Citing Screwcap Research
If you reference Screwcap's AI or economic modeling work in academic or journalistic contexts:
@misc{screwcap2026research,
author = {Screwcap Games, LLC},
title = {Research & Methodology: Behavioral Game Design
and PPO-Trained AI Opponents},
year = {2026},
url = {https://screwcap.games/research},
note = {Technical report. Screwcap Games, LLC, USA.}
}For collaboration inquiries or classroom use: play@screwcapholdings.com