Library / Mathematical Modeling DPF
Jump to passage
In this reading

Link to current text

Published source confirmed at last check

Source changed 2026-10-03 14:36:52 UTC · snapshot created 2026-10-03 14:38:14 UTC · last check 2026-10-03 15:10:10 UTC

MMP.7:5 - Archetypal Grounding

MMP.7:5.1 - Infer a success rate from a selectively submitted log

A team asks what fraction of attempts succeed. In a proposed model, each attempt succeeds with probability p. Every success is logged; each failure is logged independently with probability 1/4. Initially the team has a fixed-size sample of independently drawn log entries, with no information about how many attempts produced the source log.

The event variable Y is success or failure. The recording flag R says whether the attempt enters the log. Their joint probabilities are:

OutcomeProbability
Success, loggedp
Failure, logged(1-p)/4
Failure, unlogged3(1-p)/4

The included-case success probability is q = p / [p + (1-p)/4] = 4p/(1+3p). If the observed fraction of successes is 1/2, the likelihood estimate of q is 1/2, and transforming it gives p_hat = q_hat/(4-3q_hat) = 1/5. Sampling uncertainty remains; this calculation corrects which rate is being estimated.

The first useful result is the distinction between a 50% rate among reports and the estimated 20% rate among attempts under the supplied reporting assumptions. If the failure-reporting probability is unknown, several combinations of that probability and p can produce the same q. The log alone then leaves the population rate unresolved.

Now the procedure changes: a register names a fixed cohort of N attempts and links each submitted report to its attempt. Model those attempts as independent, each with the same success probability p, retaining the stated reporting rule. Since every success is reported, an unreported attempt is a failure. If there are k success reports, the likelihood for p is proportional to p^k (1-p)^(N-k); the failure-reporting factors do not depend on p. The estimate becomes k/N. Conditioning only on reported entries would throw away information the revised procedure provides.

This is a change in the team’s observing method. ME can describe the linked-attempt register and responsibility for recording it. Whether to introduce it depends on what resolving the population rate would change in the team’s work.

MMP.7:5.2 - Preserve a common influence across readings

Two sensors measure quantities x1 and x2 with one shared calibration offset b. Their readings are Y1=x1+b+E1 and Y2=x2+b+E2, with independent zero-mean errors of variance sigma squared. Begin by retaining b as a common parameter. The joint conditional density factors given b; each factor uses that same value.

For the difference, Y1-Y2=x1-x2+E1-E2: the offset cancels. Its error variance is 2 sigma^2. A measurement of the difference can therefore be useful while either absolute value remains uncertain.

For repeated measurements of one x, suppose an additional justified model describes the common offset as a zero-mean random variable B with variance tau squared, independent of the errors. The average of n readings has variance tau^2 + sigma^2/n. Integrating a separate offset for every reading would incorrectly produce (tau^2+sigma^2)/n. Repetition reduces independent noise but leaves this common calibration contribution.

If the instrument is independently recalibrated before every reading, the arrangement changes. A separate-offset model can then be appropriate. The governing operation is to trace which influences are shared and preserve that sharing in the probability construction.

MMP.7:5.3 - Use a timed-out trial as an interval report

A test asks how long an event takes. Each independent trial is observed until its event or a fixed timeout c. Record both V=min(T,c) and a flag D indicating whether the event occurred before timeout. As an illustrative subject assumption, let T have exponential density lambda exp(-lambda t) for t at least zero, with lambda positive.

An event at time t before c contributes the density lambda exp(-lambda t). A timeout contributes P(T>=c)=exp(-lambda c), obtained by integrating the density over the unobserved tail. For m completed trials and total observed time S, including the timeout durations, the likelihood is proportional to lambda^m exp(-lambda S).

With completions at times 1 and 2 and one timeout at 4, S=7 and m=2. Maximizing this illustrative likelihood gives lambda_hat=2/7. Treating the timeout as a third event instead gives 3/7; dropping it gives 2/3. The flag determines which operation is correct.

If only completed trials enter a database and neither the number nor identities of timed-out trials are available, the observed-time density instead conditions on completion: divide the event density by 1-exp(-lambda c) on the interval before c. A changed recording rule changes the model even when the stored times look the same. The exponential assumption is dispensable: a different duration law supplies its own event density and tail probability.