← Systems
Case Study · System 002

Sky Chess Lab.
A partner across the board.

A research through design study of an AI chess partner that shares a physical board with you instead of moving you to a screen.

Contents
  1. 00Field Note
  2. 01Research Question
  3. 02Interaction Thesis
  4. 03Intended Experience
  5. 04System Architecture
  6. 05Experiments
  7. 06What Failed
  8. 07What We Learned
  9. 08Current Limits
  10. 09Next Experiment
  11. 10Research Ethics
  12. 11Artifact Record
00 · Field Note

A partner who sits at the board.

The starting image was simple. A physical chessboard on a table, wooden pieces, and an AI partner named Sky who is present in the room rather than waiting inside an app. You move a piece with your hand. She sees it, thinks about it, and tells you her reply out loud. You move her piece for her.

Everything in this project follows from refusing the obvious shortcut. The obvious shortcut is to render a board on the screen and let the player click. That is a solved product. It is also a different experience, because the moment a digital board exists the physical one becomes decoration.

The screen should hold the partner, not the game.
01 · Research Question

Can the screen feel co-present without replacing the board?

The question under study is whether an interface can carry the sense of a person sitting across from you while the game itself stays entirely in the physical world.

Co-presence here is not a graphics problem. It is a question of attention, timing, and voice. If the player keeps looking down at the wooden board and only glances up to hear Sky respond, the design is working. If the player starts watching the screen to know the game state, the design has failed and the physical board has been demoted.

02 · Interaction Thesis

The interface contains a person, not a game.

The screen shows Sky at her desk, the conversation between you, her move instructions in plain language, and the current status of the system. That is the entire surface.

It deliberately contains no playable digital board. No coordinates grid to click, no drag targets, no mirrored position display. This is a constraint, not an omission, and it is the constraint that makes the rest of the research hard. Without a digital board there is no fallback input path, which means perception has to carry the whole system.

No digital board means perception cannot be optional.
03 · Intended Experience

Camera first, from setup to farewell.

The intended session opens with the camera, not with a menu. The player points a webcam or phone at the board and is asked to click the four corners of the playing surface so the system knows where the board lives in the frame.

Orientation is then verified so the system knows which side is which, and the player chooses a color. From there the loop is turn taking. The player moves, confirms the turn, and Sky speaks her reply through browser speech. Her move is stated in words for the player to execute by hand.

The camera lifecycle is explicit throughout. The player can pause the camera mid game and end it cleanly at the close of a session, and the interface always states whether the camera is currently watching.

04 · System Architecture

Thin server, heavy client, bounded engine.

Capture and vision run in the browser using media APIs and Canvas, so frames stay on the player's device and never travel to a server.

A Flask JSON API carries game state between the client and the rules layer. python-chess holds legality and position, and Stockfish supplies the opponent's play. Browser speech synthesis gives Sky her voice.

The division of labor is intentional. The chess side of the system is deterministic and well understood. All of the genuine uncertainty sits in the client side perception layer.

FIG 001 — Session pipeline, from camera frame to spoken replyDSL
Camera frame · browser capture, client side only
→
Board mapping · four corner homography to 64 squares
→
Square features · RGB, edge, and texture per square
→
Change detection · temporal baseline with exposure compensation
→
Move inference · candidate changes scored against legal moves
→
Rules and engine · python-chess state, Stockfish reply
→
Spoken instruction · browser speech, player moves by hand
05 · Experiments

What we actually built and tested.

Manual homography from four clicked corners, mapping the board plane to a fixed 64 square grid regardless of camera angle.

Per square feature extraction combining RGB statistics, edge density, and texture measures, on the theory that occupancy leaves a stronger signal than color alone.

Temporal baselines that hold a reference state of each square between turns, so a move appears as a difference rather than an absolute classification.

Global exposure compensation to subtract whole frame brightness shifts before looking for local change, and legal move scoring that ranks detected square changes against the set of moves the rules layer says are possible.

06 · What Failed

Two failure modes, both disqualifying.

The system produced false changes while the board sat idle, and it missed real moves while they were being made. Either failure alone breaks the experience, because a chess partner that hallucinates moves is worse than no partner.

F1

Auto exposure

Webcam auto exposure and auto white balance shifted square values with no physical change on the board, and global compensation only partly absorbed it.

F2

Shadows

Player shadows and changing room light altered per square appearance more strongly than some genuine piece movements did.

F3

Hand occlusion

A hand over the board during a move corrupted the baseline and left the system reasoning about a frame it should have discarded.

F4

Corner click error

Small inaccuracies in the four clicked corners propagated across the homography, so edge squares sampled partly off their true area.

F5

Visual ambiguity

Similar looking pieces, and light pieces on light squares, produced differences too small to separate from noise.

F6

Device variability

Thresholds tuned on one camera and one room did not transfer to another, which means the heuristics were fitting a setup rather than a task.

07 · What We Learned

Legality cannot repair perception.

The most useful finding is a negative one. Constraining detected changes to legal moves felt like it should rescue a weak signal, and it does not. When perception is noisy, legality simply selects a plausible wrong move instead of an implausible one, which makes the error harder to notice rather than less frequent.

The second finding is about sequencing. We were tuning heuristics before we had data, which meant every improvement was measured against a handful of live sessions rather than a fixed set of labeled frames. The correct order is a diagnostic dataset first, then an explicit occupancy classifier, then any further heuristic work.

Legal move filtering does not fix weak perception. It disguises it.
08 · Current Limits

Live move recognition is not dependable yet.

Stated plainly, monocular live move recognition in this system does not work reliably enough to trust in an unsupervised game. It succeeds in controlled light with a fixed camera and a cooperative player, and it degrades quickly outside those conditions.

The public repository is an experimental research prototype. It is not a product, it is not a finished chess application, and it should not be read as one. The chess and conversation layers are sound. The perception layer is an open problem under active study.

09 · Next Experiment

Data before heuristics.

The next phase replaces tuning with measurement. A diagnostic capture tool records labeled frames across cameras, lighting conditions, and board sets, producing a dataset that any change can be evaluated against.

On top of that dataset the plan is an occupied versus empty classifier per square, a separate piece color classifier, and a hand present segmentation step that suppresses inference during occlusion instead of reasoning through it.

Move inference then becomes confidence aware. Below a threshold the system says it is unsure rather than guessing, and a recovery path lets the player restate the position and continue without abandoning the game.

10 · Research Ethics

A demo is not a detector.

Any scripted concept demonstration of this system must be labeled as scripted. Showing the intended experience is legitimate research communication. Showing it in a way that implies the camera detector is complete is not.

This distinction matters more than usual here, because the failure is invisible to an audience. A scripted run and a working run look identical on video, which places the burden of honesty entirely on the label.

11 · Artifact Record

What is published.

Author: Fatima Aguilar, Decision Systems Lab.

Public repository containing the client side vision pipeline, the Flask JSON API, the python-chess and Stockfish integration, and the conversational interface, published in its current experimental state.

A six page technical paper documenting the architecture, the experiments, the failure analysis, and the planned next phase, available below.

Every system in the Decision Systems Lab is engineered to outlive its author : a small, durable piece of operational thinking made legible to the next engineer.

Technical Report

The complete systems engineering case study, including data flow specifications, rule definitions, validation harness, and deployment notes, is available as a downloadable PDF.