← Back to Games

Research & Methodology

The Science
Behind the Games

Behavioral economics. Reinforcement learning. DSGE modeling. Cognitive psychology. UI research. These are not decorative influences — they are architectural constraints.

Behavioral EconomicsPPO / Reinforcement LearningDSGE ModelingCognitive Load Theory

01

Abstract

This document describes the methodological foundations underlying Screwcap Games, LLC's portfolio of browser-based games. Our design philosophy is predicated on the proposition that entertainment and education are not in tension — that genuinely calibrated challenge, grounded in real economic and behavioral models, produces both higher engagement and measurable transfer of skill.

The Screwcap stack draws from five disciplines: reinforcement learning (DoubleFives PPO agent), DSGE modeling (www.chairmanmode.app), behavioral economics (Gold Digger), cognitive psychology (difficulty scaling), and UI/UX research (interface friction). These are not marketing claims. They are architectural constraints embedded in how each game is built.

02

AI Methodology: DoubleFives PPO Agent

The AI powering DoubleFives is a Proximal Policy Optimization (PPO) agent trained via deep reinforcement learning in a four-player adversarial domino environment. PPO belongs to the same algorithm family as OpenAI Five and the predecessor approaches to AlphaGo — it is among the most battle-tested policy gradient methods in the field.

The agent uses an Actor-Critic architecture with a 3-layer MLP policy network. Seven distinct AI personalities were developed by training separate policy heads with modified reward functions — introducing different weightings on aggressive play, defensive blocking, and partner cooperation. These produce genuinely distinct behavioral profiles, not cosmetic variation.

Model weights are exported via ONNXand run entirely in-browser. No game moves are transmitted to a server. The AI runs locally on the player's device.

03

Economic Model Design: www.chairmanmode.app

www.chairmanmode.app is built on a Dynamic Stochastic General Equilibrium (DSGE) model — the workhorse framework of modern monetary policy analysis, used by the Federal Reserve, ECB, and major academic institutions. It implements multiple historically calibrated scenarios — 1929, 2008, and the Volcker Shock among them — with parameters drawn from FRED.

It implements the Taylor Rule as a policy benchmark:

it = r* + π* + 1.5(πt − π*) + 0.5(yt − ȳt)

04

Behavioral Economics in Game Design

Gold Digger's prediction market framing exploits the well-documented gap between objective probability and subjective probability weighting identified by Kahneman and Tversky. Players systematically overweight small probabilities — the game makes this miscalibration visible through immediate feedback. The pedagogical goal is probability calibration.

DoubleFives' difficulty design is built the same way, one layer down: seven policy heads, each a separate PPO training run with a different reward weighting on aggressive scoring, defensive blocking, and partner cooperation, give seven genuinely distinct opponents rather than one AI throttled by a difficulty slider. That range is what lets the game sit each player at the boundary between anxiety and boredom that Csikszentmihalyi's Flow theory describes, across two, three, and four-seat play, rather than tuning a single opponent down until it stops being a real game.

05

Rating Methodology

DoubleFives agents are rated with Glicko, not classic Elo. Elo assumes every rating is equally certain; Glicko carries a rating deviation alongside the rating itself, which matters here because a new policy head plays thousands of games against the established roster within days — treating those early results as exactly as reliable as a rating with a year of games behind it would be the wrong kind of confident.

g(RD) = 1 / √(1 + 3q²RD²/π²)

As of the latest measured checkpoint, the strongest neural policy head rates meaningfully above the scripted baselines it trains against (greedy-scoring, tile-counting, and random-play bots) — the gap the game means when it says the AI “counted better.” The board is public and updates as training continues; see the live figures on the homepage Lab section.

06

UI/UX Design Principles

Screwcap interfaces minimize extraneous cognitive load while preserving germane cognitive load — the effortful engagement required to learn the underlying model. A player struggling with TheChair's rate decisions should be struggling with the macroeconomics, not the UI.

07

Raw Training Metrics

Live from the same feed as the homepage Lab section — not a snapshot pinned to one training run.

● ● ●
// DoubleFives PPO — Training Summary
algorithmProximal Policy Optimization
total_games_playedloading…
top_rated_agentloading…
gpuNVIDIA RTX 3090 (24GB VRAM)
frameworkPyTorch 2.x + Stable-Baselines3
architectureActor-Critic, 3-layer MLP
personalities7 distinct policy heads
deployment_formatONNX (client-side, no latency)
statusTraining in progress

08

Citing Screwcap Research

If you reference Screwcap's AI or economic modeling work in academic or journalistic contexts:

@misc{screwcap2026research,
  author    = {Screwcap Games, LLC},
  title     = {Research & Methodology: Behavioral Game Design
               and PPO-Trained AI Opponents},
  year      = {2026},
  url       = {https://screwcap.games/research},
  note      = {Technical report. Screwcap Games, LLC, USA.}
}

For collaboration inquiries or classroom use: play@screwcapholdings.com