Skip to contentOpen Research Lab

Grouped by research area. Each entry says what can be claimed right now and what kind of evidence supports it. 日本語

Notes are grouped by what they contribute, not by how confident they sound. A Finding reports an observed result, a Method defines a research interface, and a Protocol freezes a future test before confirmatory data exists.

GPU cluster scheduling

Question: Under what conditions does size-based scheduling — running the shortest job first — break?

Current state: The three quantities to watch are largest-job share of pool capacity × concurrent job count × large-job frequency. In the 256-server, rho 0.85 synthetic model, the rule-of-thumb ratios are 0.92 (frequency 0.002) or 0.72 (frequency 0.02) at concurrency 10, and 0.73 or 0.66 at concurrency 30. Dropping frequency makes the number conservative for rare large jobs and dangerous for frequent ones.

In a new cluster, inspect whether mean response time grows with the observation window rather than using completion rate. If doubling the window raises the mean by at least 1.4×, that class is diverging. The real-cluster reach is extremely narrow: ten of Philly’s 11 virtual clusters contain no job above a quarter of capacity, and the sole exception has exactly one job out of 19,100 above half of capacity.

Evidence boundary: This is not a scheduler comparison on a real arrival sequence, and there are no independent external replications.

Read the current Research Note (EP-0013) → — Finding

See the current state and revision history →

Queueing system identification

Question: Given a queueing world whose mechanism is hidden, how much of it can external records alone recover?

Current state: Within a disclosed family of candidate mechanisms, one hidden instance was identified correctly and all five structural components matched the withheld truth.

Evidence boundary: This tests selection from given options, not discovery from an open hypothesis space, and transfer to a real system has not been tested.

Read the current Research Note (EP-0001) → — Finding

See the current state and revision history →

Human movement model interfaces

Question: Before measured human motion reaches a model, can explicit contracts and dataset boundaries stop coordinate mix-ups, target leakage, and false external validation?

Current state: The interface contract rejects seven deliberately broken variants. The dataset gate splits Carter into 30 development, 10 validation, and 10 test participants; reserves OpenCap laboratory data for external validation; and reserves Knee Grand Challenge for internal-load stress testing. B3D is a source-adapter format, the internal representation is NumPy plus explicit semantics, and Nimble is not mandatory.

Evidence boundary: No model has been fitted and no holdout has been read.

Read the current Research Note (EP-0005) → — Dataset

See the current state and revision history →

Badminton biomechanics

Question: In the smash, can the effect of short preparation time be separated from the effect of backward centre of mass?

Current state: Only the experimental design and power sensitivity are public; no observations exist. Detecting the interaction requires substantially more participants than detecting the main effects, so the sample size remains undecided.

Evidence boundary: Instrument feasibility and ethics remain unresolved, so the protocol is not sealed and data collection is not authorized.

Read the current Research Note (EP-0008) → — Protocol

See the current state and revision history →

The instrument in domains that require acting

Question: What can a seal-and-score apparatus measure when the predictor is also the actor, and where does the observational contract break?

Current state: The observational prediction contract breaks in five places: a miss cannot be split into a wrong model and nobody acting, the predictor can go and make its own prediction false, the scoring date is not theirs to choose, waiting becomes a stall, and success moves the distribution being measured.

Each breach is answered by an invariant that refuses the record. Scoring runs only through a four-point joint — threshold, decision, execution, result — and a prediction whose decision never opened can be neither supported nor refuted. Self-interference is settled by a stance declared at seal time; counting an intervention as its own success was declined, because it would stop the ledger separating a good model from a good operator.

Evidence boundary: Contract validation only. No prediction sealed in this grammar has reached a scoring date, and there is no guarantee that five is the complete set of breaches.

Read the current Research Note (EP-0001) → — Method

See the current state and revision history →

What a measurement design sees, and what it does not

Question: Where does a measurement-design result obtained on a world of your own making give way under a sweep?

Current state: Both headline results are currently withdrawn. A synthetic world flatters its author, and the flattery turned out to have forms.

The first was fixing the only thing that drives the quantity being measured. That constant moves the metric by 0.209; the thing we claimed to measure moves it by at most 0.027, with an unstable sign — a factor of 7.9. The second was pinning an unobservable nuisance at one value and publishing a decision line as a function of sample size. Sweep the nuisance and a frontier that fully closes clears the published line.

Only the weaker form survives: reuse and the settling rate are different quantities, and only reuse responds. That one is not an artifact of a constant.

Evidence boundary: Synthetic, and the mechanism family is ours. No real data of any kind was used and no claim about any application field is made. This is a self-review, so an attack we did not think of is by construction not in it.

Read the current Research Note (FRONTIER-METRICS-REVIEW-0001) → — Negative Result

See the current state and revision history →

Browse all published Research Notes by date →


Future records may include Replication, Negative Result, Dataset, and Benchmark without changing existing research identities.