banner-128m.png

ShoreCrab-128M: Scoring Answers Instead of Generating Them — a 26-ms Multiple-Choice Visual Reasoner

Haverbex · Technical report v0.1 · Research preview · Updated 4 October 2026

<aside> 🤝

I am a student who runs a company on my own. Investment is welcome to support further research and the follow-up work to this project.

저는 혼자 일하며 회사를 운영하는 학생입니다. 다양한 연구와 이 연구의 후속 연구를 위해 투자를 환영합니다.

contact: [email protected]

</aside>

Abstract

ShoreCrab maps an image, a question and a supplied answer set to a joint distribution over the candidates and “none of these.” The selected I-A60 checkpoint has 128.1M parameters and a median decision latency of 25.6 ms on an Apple M5 Pro GPU and 109.3 ms on CPU. It reaches 70.77% on a 2,600-question GQA-derived development set; three final-stage seed runs average 70.50%, with a sample standard deviation of 0.68 percentage points. On a matched 1,000-question development subset, accuracy is 70.0%, compared with 47.1% for zero-shot SmolVLM-256M. These development measurements do not establish a general VLM ranking. A CLI game controller uses A60 visual features with a separate game head; rendered everyday and robotics examples use the original candidate reader and include failures.

1. Problem setting

The input is an image I, a question q and K text candidates, with 2 ≤ K ≤ 16. The candidates are part of the request and can change without changing the output vocabulary. The model returns K candidate probabilities and one probability for an unresolved answer.

This interface targets bounded visual decisions: selecting an action from a game screen, identifying an item from a catalogue or choosing an object in a camera view. The output is a distribution over supplied alternatives; it does not generate an open-ended explanation.

2. Method

overview.png

Figure 1. Image tokens pass through a question-conditioned resampler and a shared candidate reader to produce joint candidate/none probabilities. The displayed probabilities illustrate the interface. Section 2.1 specifies the selected checkpoint’s view and tiling configuration.

2.1 Visual reading and question conditioning

The selected configuration uses a 384-dimensional representation, six attention heads and two shared reader layers. An image fitting the 512-pixel canvas is read in one view. Larger inputs add native-resolution local tiles with 480-pixel cores and a 16-pixel halo. The overview uses 64 resampler queries, and each local view uses 32.

A learned query bank is shifted by the question embedding before attending to image features. Visual compression is therefore conditioned on the question. The learned bank supplies the residual so that the queries remain distinct.

2.2 Shared candidate reader

Each candidate’s text tokens attend to the same compressed visual tokens and question tokens. Reader weights and the score projection are shared across candidates. All candidate scores are produced in one forward pass.

$$ Z = \mathrm{Resampler}(E_{\mathrm{img}}(I), E_{\mathrm{txt}}(q)), \qquad h_k = \mathrm{Reader}(E_{\mathrm{txt}}(c_k), [Z; E_{\mathrm{txt}}(q)]), \qquad s_k = w^\top h_k . $$

2.3 Answerability and joint output