Universal Textual Teaching for LLMs

Zhanyi Lu1, Huan Wang1,*

1Westlake University

* Corresponding author: wanghuan@westlake.edu.cn

Westlake University
UTT distills multi-role teaching interactions into a textual Primer. Task-specific Primers transfer to Qwen3.6-27B, yielding gains up to 41.8 percentage points on KernelBench and 31.5 on Omni-MATH-2.
Universal Textual Teaching (UTT) distills an LLM’s knowledge into a textual Primer through multi-role interactions. The Primer improves its source Student and transfers to other Students without resynthesis. When Qwen3.6-27B does not participate in synthesis, accuracy gains reach 41.8 and 31.5 percentage points on KernelBench and Omni-MATH-2, respectively.

Abstract

Knowledge distillation (KD) transfers knowledge from stronger Teacher models to weaker Student models, but most methods require training the Student parameters, thereby binding the distilled knowledge to a specific architecture and checkpoint. This implicit representation is difficult to interpret or reuse across models and limits KD for API-only or costly-to-train models. This paper studies knowledge transfer for large language models (LLMs). We introduce Universal Textual Teaching (UTT), a parameter-update-free framework that distills observed Teacher–Student knowledge gaps into a textual, interpretable, and reusable natural-language artifact called Primer.

Specifically, UTT first identifies representative gap cases through paired evaluations, and iteratively updates the Primer via multi-role interactions: the Student attempts each task, the Prompter turns evaluation feedback into a teaching instruction, the Teacher provides a targeted demonstration, and the Synthesizer consolidates validated lessons. Empirically, on the challenging math (Omni-MATH-2) and code generation (KernelBench) tasks, extensive results confirm the effectiveness of the method: UTT remarkably raises the Student’s accuracy from 9.4% to 48.6% and Fast₁ accuracy from 9% to 35% on KernelBench, while increasing mathematical reasoning accuracy from 27.6% to 51.7%. UTT also performs better than representative prompt engineering and parameter-based KD methods. Of note, UTT is shown to be generalizable across different Teachers and Students: a Primer synthesized for one Teacher–Student pair can generalize to other Students that do not participate in the synthesis.

Overview of our UTT method

Paired Teacher and Student evaluation identifies knowledge gaps. Student attempts, Prompter instructions, Teacher demonstrations and Synthesizer updates create a Primer for source and transferred Students.
Illustration of the proposed UTT framework. Paired evaluation identifies Teacher–Student knowledge gaps. The Student attempts each task; the Prompter diagnoses feedback; the Teacher provides a demonstration; and the Synthesizer consolidates validated teaching records. Candidate Primers are accepted when accuracy on the current batch does not decrease. All model parameters remain fixed.

The Teacher model also serves as the Prompter and Synthesizer. The final Primer is prepended to task inputs and can be reused by another Student without resynthesis.

Main Results

Models. Pro: DeepSeek V4 Pro; Flash: DeepSeek V4 Flash; Qwen: Qwen3.6-27B; Opus: Claude Opus 5.

More Results and Analysis

Primer Examples

Full examples of the generated Primers for mathematical reasoning, GPU kernel generation, and multimodal geometry.

Mathematical reasoning · Pro → Flash
Full Primer · Omni-MATH-2
Download text

Primer: Olympiad Problem-Solving Patterns

0 Output discipline

  • End with a finite checkable proof/boxed answer; prefer compact algebra, counting, or graph arguments.
  • Verify definitions, boundary cases, and small examples before trusting analogies or known results.

1 Geometry

  • Circle tangency via coordinates: normalize; circle through points x2+y2+ax+by+c=0x^2+y^2+ax+by+c=0; prove common point and proportional gradients.
  • Billiards by unfolding: reflect polygon, not ray; path becomes straight segment in tiling; use gcd/color conditions for vertex hits.
  • Distinguish metric vs combinatorial definitions; use explicit polar/nonregular counterexamples.
  • Tetrahedron altitude from six edges: build base B′C′D′B'C'D'; build rotated face points ABA_B opposite side of C′D′C'D' with ABC′=AC,ABD′=ADA_BC'=AC, A_BD'=AD, and ACA_C opposite side of B′D′B'D' with ACB′=AB,ACD′=ADA_CB'=AB, A_CD'=AD. Perpendiculars from ABA_B to line C′D′C'D' and ACA_C to line B′D′B'D' meet at the true foot X′X'. Altitude =∣ABY′∣2−∣X′Y′∣2=\sqrt{|A_BY'|^2-|X'Y'|^2}, constructible by Pythagorean difference. Use full supporting lines, not segments. Foot from AA to line BCBC has signed BP=(AB2+BC2−AC2)/(2BC)BP=(AB^2+BC^2-AC^2)/(2BC); BP<0BP<0 or BP>BCBP>BC means foot outside segment. Transfer signed positions; unsigned distances fail in obtuse cases.

2 Number theory and polynomial values

2.1 Vieta jumping/descent

For (a+b)(a+b+1)/(ab)=N(a+b)(a+b+1)/(ab)=N: fix NN, rewrite as quadratic, Vieta gives positive integer other root; descend to equal pair; verify integrality/positivity.

2.2 Lyndon words and fractional parts

Suffix/prefix lex conditions become Lyndon words; count length LL over qq by 1L∑d∣Lμ(d)qL/d\frac1L\sum_{d\mid L}\mu(d)q^{L/d}.

2.3 Reduced denominators / averaging

For Sn=An/n!S_n=A_n/n!, denominator n!/gcd⁡(An,n!)n!/\gcd(A_n,n!); lifts Ar+pk≡Ar−pk(modpk+1)A_{r+p^k}\equiv A_r-p^k\pmod{p^{k+1}}; existence via averaging and Stirling.

2.4 Sum-of-kk-others and omitted elements

If each element is a sum of kk others in a set of k+mk+m, count omitted elements; compare largest/smallest omitted sums. Don't use sign-count bounds: largest need not be positive sum. For zero-sum symmetric sets solve a+b+c=−2xa+b+c=-2x.

2.5 Digit sums and carries

σ(10m+i)=σ(m)+i\sigma(10m+i)=\sigma(m)+i; safe block iff σ(m)≡1(mod11)\sigma(m)\equiv1\pmod{11}; trailing-9s give σ(m+1)−σ(m)=1−9t\sigma(m+1)-\sigma(m)=1-9t. Maximal safe intervals shape 9+10+10+99+10+10+9.

2.6 Integer-valued polynomials and P(i)P(i)
  • Integer-valued iff P(x)=∑ck(xk)P(x)=\sum c_k\binom{x}{k}, ck∈Zc_k\in\mathbb Z.
  • At x=ix=i, analyze vπ((ik))=vπ(Nk/k!)v_\pi(\binom{i}{k})=v_\pi(N_k/k!). For p≡1(mod4)p\equiv1\pmod4, p=ππˉp=\pi\bar\pi in Z[i]\mathbb Z[i]; Hensel gives ss with πe∣i−s\pi^e\mid i-s, so among i,…,i−k+1i,\dots,i-k+1 at least ⌊k/pe⌋\lfloor k/p^e\rfloor are divisible by πe\pi^e; hence no p≡1(mod4)p\equiv1\pmod4 appears in denominators.
  • p=2p=2 or p≡3(mod4)p\equiv3\pmod4 attainable: 1/2=6(x4)+31/2=6\binom{x}{4}+3; 1/p=(p−1)!(xp)1/p=(p-1)!\binom{x}{p}.
  • Answer: a+bia+bi, a,b∈Qa,b\in\mathbb Q, with νp(a),νp(b)≥0\nu_p(a),\nu_p(b)\ge0 for every p≡1(mod4)p\equiv1\pmod4. Not all Q(i)\mathbb Q(i): 1/51/5 excluded.
2.7 Prime-divisor constraints ω(n)>K\omega(n)>K
  • Strict >K>K means minimal ω(n)=K+1\omega(n)=K+1; constants cc need c>0c>0, ω(c)≤K+1\omega(c)\le K+1. Boundary ω(c)=K+2\omega(c)=K+2 with ω(n)=K+1\omega(n)=K+1.
  • Monomials xmx^m work.
  • Nonconstant nonmonomials fail. If P(0)≠0P(0)\ne0, Schur gives K+2K+2 primes qiq_i with nonzero roots ai mod qia_i\bmod q_i; CRT+Dirichlet builds nn with exactly K+1K+1 prime factors and n≡ai mod qin\equiv a_i\bmod q_i; then qi∣P(n)q_i\mid P(n), qi∤nq_i\nmid n, so ω(P(n))≥K+2\omega(P(n))\ge K+2 or P(n)≤0P(n)\le0. If P(0)=0P(0)=0, factor xmQ(x)x^mQ(x), Q(0)≠0Q(0)\ne0, apply to QQ; handle constant factor c>1c>1.

3 Scheduling and exact-degree constructions

  • Tournament stays: interval per player; pairwise intersection + Helly gives common day; one match at central day; total cost baseline 2M2M plus idle player-days; balance distinct players before/after.
  • Touch/degree: for nn objects each touching exactly 3 others, handshake gives 3n3n even, so nn even; do not infer divisibility by 4. For large even nn, cycle base (outside squares on 4m4m-gon plus reflected pairs gives 6m6m) and insert four-square gadgets at cuts to add 8 or 16, covering residues mod 6; verify no unintended intersections.

4 Hamiltonian paths on grids

  • Numbering an n×nn\times n grid is a Hamiltonian path; diagonal projection gives a ±1\pm1-walk, main diagonal visits same parity. Bound first/last visits using side color counts; construct explicitly.

5 Permutation reachability via swap graphs

  • Allowed swaps = graph edges; connected graph iff every permutation reachable. Prove connectivity by explicit spanning path; disprove by separated component.

6 Local block constraints and two-stage counting

  • Count distinguished positions first, then assign remaining; encode as 0,±10,\pm1. For 2×22\times2 zero-sum blocks, general solution aij=(−1)i+j(ri+cj)a_{ij}=(-1)^{i+j}(r_i+c_j); propagate sign choices across overlaps.

7 Hyperplane-generic finite sets

  • Minimal kk-generic sets: lower via concurrent lines with kk points; upper via private-point subcover and affine linear functions; incidence/basis gives ∣M∣≤kn|M|\le kn for k,n>1k,n>1.

8 Partition minima and smoothing

  • Define S(N)=min⁡n1+⋯+nm=N∑f(ni)S(N)=\min_{n_1+\cdots+n_m=N}\sum f(n_i) for nonincreasing ff. Balanced partition is only an upper bound: S(N)≤rf(q+1)+(m−r)f(q)S(N)\le r f(q+1)+(m-r)f(q), N=mq+rN=mq+r, 0≤r<m0\le r<m. Summing one fixed partition family overcounts because SS is the minimum over all partitions. Count all zero parts: for m=20m=20, N=0,…,19N=0,\dots,19, coefficient of f(0)f(0) is 20+⋯+1=21020+\cdots+1=210, not 20.
  • Smoothing extremal: if ∑f(i)=A=(a+12)\sum f(i)=A=\binom{a+1}{2}, then ∑N≥0S(N)≤∑N=0ma(ma−N)=ma(ma+1)/2\sum_{N\ge0}S(N)\le \sum_{N=0}^{ma}(ma-N)=ma(ma+1)/2. Equality at f(i)=max⁡(a−i,0)f(i)=\max(a-i,0), g(N)=max⁡(ma−N,0)g(N)=\max(ma-N,0). Verify: f(ni)≥a−nif(n_i)\ge a-n_i, so sum ≥ma−N\ge ma-N, and sum ≥0\ge0.

9 Fair bounded selection with biased coins

  • Do not assume arbitrary outcome probabilities can be partitioned into nn equal groups. For n=3n=3, p=1/3p=1/3, L=3L=3, weights 8,4,4,4,2,2,2,18,4,4,4,2,2,2,1 cannot split into three sums 9 (8 must pair with 1; remaining even weights cannot sum 9).
  • Verified construction: choose p1=1/np_1=1/n, p2=1/2p_2=1/2; take cc with N=2c≥n(n−1)N=2^c\ge n(n-1), announce L=c+1L=c+1. Toss coin1 once, coin2 cc times. In units 1/(nN)1/(nN), there are NN TT-outcomes of weight n−1n-1 and NN HH-outcomes of weight 1; each person needs total weight NN. Write N=qn+rN=qn+r, 0≤r<n0\le r<n; set ai=q+1a_i=q+1 for rr people, ai=qa_i=q otherwise; bi=N−ai(n−1)b_i=N-a_i(n-1). Feasibility ai(n−1)≤Na_i(n-1)\le N follows from N≥n(n−1)N\ge n(n-1). Assign aia_i TT-strings and bib_i HH-strings to person ii; probability ai(n−1)/(nN)+bi/(nN)=1/na_i(n-1)/(nN)+b_i/(nN)=1/n. Explicitly verify the partition before announcing.

10 Validation checklist

  • Integer-valued polynomials: binomial basis; check split primes p≡1(mod4)p\equiv1\pmod4 separately.
  • ω(n)>K\omega(n)>K: convert to K+1K+1; constants separately; Schur+CRT/Dirichlet to force prime divisors.
  • Touch/degree: handshake parity; avoid extra divisibility; test constructions for unintended coincidences.
  • Scheduling/grid/digit/local-counting: include idle days; diagonal parity and explicit attainment; carry blocks; propagate overlap consistency.
  • Vieta jumping: verify positive integral new root. Circle tangency: check membership and gradients. Counting formulas: check divisors/Möbius signs.
  • Tetrahedron altitude: use supporting lines and signed foot distances; test obtuse cases (BP<0BP<0 or BP>BCBP>BC); verify intersection is true projection.
  • Partition minima: balanced partition is upper bound not exact; count zero coefficients; verify extremal construction.
  • Fair selection: verify explicit equiprobable partition for announced bound; test small nn/parity counterexamples.

GPU kernel generation · Pro → Flash
Full Primer · KernelBench
Download text

PRIMER: Writing Custom CUDA Operators for PyTorch Models

Core Requirements

  • Replace operators with a custom CUDA implementation, typically inside a torch.autograd.Function or as a direct extension function.
  • No try/except or fallback logic — let assertions crash.
  • Output format is critical: the entire answer must be raw Python source, starting with the first character and continuing until the end. If the instruction says “output only the code,” the response must be solely the code block—no introductory or trailing text and no Markdown fences.
  • All methods must be fully implemented, with no placeholders.

Token Budget & Code-First Strategy

  • Code output must come immediately. Spending the budget on reasoning before the code can lead to truncation, leaving no answer at all.
  • Prefer a minimal, correct kernel, such as one thread per output element or a serial scan, to remain within the token limit.
  • For inference-only tasks, skip backward computation by raising NotImplementedError.
  • Target roughly 60–80 lines of model code; simplify kernels that grow substantially larger.

Symbol Visibility Before Binding (Critical)

Every function referenced by m.def must be known at the point of registration. The extension binding code is a plain C++ translation unit and does not permit linking later to unresolved symbols. Two safe patterns are:

  1. Monolithic source (recommended): define all functions before the PYBIND11_MODULE block inside a single CUDA source string and leave cpp_sources=[].
  2. Multi-file: declare the wrapper function in a header included by the binding file. Its definition in a separate .cu file must exactly match the declaration.

The host function signature must be compatible with torch::wrap_pybind_function, conceptually a std::function accepting the specified argument types and returning a torch::Tensor. Using &my_func directly is simpler and preferred.

Monolithic Template (Safe)

cuda_src = r'''
#include <torch/extension.h>
__global__ void my_kernel(...) { ... }
void my_op(torch::Tensor x, torch::Tensor y) { ... /* launch kernel */ }
PYBIND11_MODULE(TORCH_EXTENSION_NAME, m) {
  m.def("my_op", &my_op, "doc");
}
'''
ext = load_inline(name="my_ext", cpp_sources=[],
                  cuda_sources=[cuda_src], verbose=False)

General Kernel Design Patterns

  • One thread per output for convolution, pooling, and similar operations: launch one thread per output element, loop over kernel and channel dimensions, and use __ldg for read-only data.
  • Scan-type operations such as cumulative sums: default to one thread per independent slice, looping sequentially over the scan dimension. This is simple, token-efficient, and reliable.
  • Reverse cumulative sums: loop from right to left rather than composing flip, cumulative sum, and another flip.
  • Multidimensional indexing: parenthesize expressions aggressively to avoid compilation failures caused by missing parentheses.

Implementing Scan Operations

  • Treat the tensor as (num_vectors, L), moving the target dimension to the final position when necessary.
  • Launch num_vectors threads, each performing a serial inclusive scan.
  • For a reverse cumulative sum, iterate from right to left.
  • Let the host wrapper move the target dimension to the end, launch on a two-dimensional view, and permute the result back.

Extension Building with load_inline

  • Exactly one PYBIND11_MODULE must appear, with TORCH_EXTENSION_NAME as its first argument.
  • The name supplied to load_inline must match the module name represented by the macro.
  • Use extra_cflags=["-O3"] and extra_cuda_cflags=["-O3"].
  • Compile once and cache the module at the class level to avoid recompilation for every instance.

Lazy Compilation (Class-Level Caching)

class ModelNew(nn.Module):
    _ext = None
    def __init__(self, dim):
        if ModelNew._ext is None:
            ModelNew._ext = load_inline(
                name="op",
                cpp_sources=[],
                cuda_sources=[cuda_src],
                verbose=False)
        self.ext = ModelNew._ext

Testing Builds Locally

Verify that a small extension compiles and executes before integrating it:

ext = load_inline(name="test", cpp_sources=[],
                  cuda_sources=[cuda_src], verbose=True)
x = torch.randn(1, 1, 8, 8, 8).cuda()
ext.my_op(x, ...)

Common Failure Modes

Symptom Root cause and fix
SyntaxError at the start of the compiled stringMarkdown fences or leading text were included; output pure Python source.
No code output; token budget exhaustedThe model spent the budget on reasoning; emit code immediately.
redefinition of PyInit_xxxMultiple PYBIND11_MODULE blocks; retain exactly one.
my_func was not declaredThe function is not visible before m.def; define it first or include its declaration.
TypeError: SupportsFloatThe host function expects raw pointers; accept torch::Tensor and extract pointers internally.
Wrong output or segmentation faultNon-contiguous strides were ignored; call .contiguous() or handle strides explicitly.
CUDA expected a semicolonA complex index expression has unbalanced parentheses; introduce intermediate variables and balance the expression.
Extension-name mismatchThe supplied name and binding module differ; use PYBIND11_MODULE(TORCH_EXTENSION_NAME, ...).

Verification Checklist

  • Output is pure Python source, without fences or commentary when these are forbidden.
  • Code is emitted immediately without prolonged reasoning.
  • Exactly one PYBIND11_MODULE is present, and every referenced host function is visible beforehand.
  • Host functions accept PyTorch-compatible types rather than raw pointers.
  • No fallback logic appears in the optimized implementation.
  • Tensor contiguity is checked or handled.
  • Reverse scans traverse in reverse rather than composing flips.
  • Multidimensional index expressions are balanced and tested.
  • The extension is tested on a small tensor before a full benchmark run.

Multimodal geometry · Opus → Qwen
Full Primer · MathVerse
Download text

PRIMER: Multiple-Choice Geometry from a Figure

0. Non-negotiable output rule

One letter is the deliverable. Write "Answer: X" first (or end with the bare letter on its own line), then 2–6 lines of justification. Nearly every observed failure was identical: no final output — budget burned re-reading the figure, cataloguing every labeled value, or debating which vertex owns a number. Partial reasoning scores zero; a best guess always beats silence.

Budget policy: 2–4 reasoning steps, at most one or two figure interpretations, then stop.

  • Value matches an option →\rightarrow commit; stop verifying.
  • Matches nothing →\rightarrow say so in one sentence, switch the assignment once, else pick the nearest listed value.
  • Never solve for every lettered unknown; stop the instant the target is pinned.
  • "Cannot be determined" only after a determinate chain visibly breaks — never as a hedge.

1. Read the figure, then isolate one unknown

  • Attach each number to the vertex/side/bracket its mark touches; name it ("∠MLJ=30∘\angle MLJ=30^\circ"). An arrow from a value just labels the adjacent lettered angle (r∘←90r^\circ\leftarrow90 means r=90r=90).
  • Decide what each length denotes: radius vs diameter, whole bracket vs sub-segment, side vs apothem, slant vs height. Label-swapping is the #1 distractor generator. An arrow drawn inside a small circle from its center is that circle's radius.
  • List only the structural facts you need: which point-triples are collinear, which segments are parallel/tangent.
  • Chain, don't systematize: anchor on the triangle/relation with two knowns, propagate, stop.
  • Prune decoys explicitly; extra labels often serve other sub-questions.
  • Not-to-scale figures: test candidate readings against options; keep the clean one.

2. Angle propagation toolkit

Two tools carry most crossing-line figures: vertical angles at the intersection and 180∘180^\circ sums (triangle, straight angle). Template: triangle with two given angles →\rightarrow third angle →\rightarrow vertical angle across →\rightarrow next triangle →\rightarrow straight-angle subtraction →\rightarrow target.

Parallel lines: arrowheads declare the parallel pair; find the transversal joining target to given. Corresponding/alternate interior ⇒\Rightarrow equal; co-interior or linear pair ⇒\Rightarrow supplementary. Misreading same-side as alternate gives the 180−θ180-\theta distractor. Extended ray ⇒\Rightarrow linear pair. In a parallelogram a diagonal is a transversal for both side pairs; pair the correct halves. Quadrilateral interior sum 360∘360^\circ.

3. Congruent/joined figures

Congruent polygons glued along a shared side: corresponding angles transfer; the angle at a seam vertex is usually the sum of one angle from each. Validate with polygon angle sums.

4. Similar triangles (parallels, shadows, mirrors)

AB∥CDAB\parallel CD with apex PP ⇒△PAB∼△PCD\Rightarrow\triangle PAB\sim\triangle PCD; h(P→AB)=(AB/CD)⋅h(P→CD)h(P\rightarrow AB)=(AB/CD)\cdot h(P\rightarrow CD). Distance between parallels =htotal−hsub=h_{\mathrm{total}}-h_{\mathrm{sub}} (must be << total). Mirror/shadow: hfar/hnear=dfar/dnearh_{\mathrm{far}}/h_{\mathrm{near}} =d_{\mathrm{far}}/d_{\mathrm{near}}.

5. Circle toolkit

  • Find diameters first: collinear labeled points through the center ⇒\Rightarrow adjacent central angles are a linear pair; all central angles sum 360.
  • Arc notation: two letters = minor arc = central angle; three letters = arc through the middle letter, usually major ⇒360−minor\Rightarrow360-\mathrm{minor} (>180∘>180^\circ).
  • Inscribed = 12\frac12 arc; central = 2×2\timesinscribed; cyclic quad opposite angles supplementary; chord = 2Rsin⁡A2R\sin A.
  • Tangent from external PP: tangent ⊥\perp radius; AP=OP⋅sin⁡∠AOPAP=OP\cdot\sin\angle AOP, r=OP⋅cos⁡∠AOPr=OP\cdot\cos\angle AOP; check r<OPr<OP.
  • Equidistant chords ⇒\Rightarrow congruent chords ⇒\Rightarrow congruent arcs ⇒\Rightarrow equal inscribed angles.
  • Internally tangent circles: centers and contact point are collinear. Two equal circles of radius rr internally tangent to a big circle at opposite ends of a diameter and tangent to each other ⇒R=2r\Rightarrow R=2r.
  • Perimeter of a curvilinear region = sum of arc lengths, each (θ/360)⋅2πρ(\theta/360)\cdot2\pi\rho with its own radius. Half of a circle of radius ρ\rho contributes πρ\pi\rho. Typical shaded lune: πR+πr+πr\pi R+\pi r+\pi r. Rectangle on a semicircle diameter: h=R2−(w/2)2h=\sqrt{R^2-(w/2)^2}.

6. Area by symmetry and cancellation

Shaded = whole −- unshaded; hunt equal-area pieces first. Common apex ⇒\Rightarrow compare triangles by base alone: apex at a rectangle's center gives every top/bottom triangle height H/2H/2, area bH/4bH/4, so group bases until they sum to a full side (top set = bottom set = 14\frac14 of area; sides = 12\frac12). If a shaded base and its mirror unshaded base are equal, shaded bases total one full width ⇒14\Rightarrow\frac14 of the rectangle. Verify all region fractions sum to 1. Assume convenient dimensions (4×24\times2) or coordinates. Quarter disc radius rr and semicircle on its chord both have area πr2/4\pi r^2/4 ⇒\Rightarrow shaded = 12r2\frac12r^2; π\pi-free answers signal cancellation.

7. Cone/net toolkit

(θ/360)⋅2πR=2πr(\theta/360)\cdot2\pi R=2\pi r; θ=90∘⇒R=4r\theta=90^\circ\Rightarrow R=4r. Lateral surface = πrl\pi rl; l2=h2+r2l^2=h^2+r^2. AE=AFAE=AF, DE=DF⇒DE=DF\Rightarrow axis ⊥\perp base.

8. One-line validation, then emit

Supplements sum 180; polygon sums close; drawn obtuse ⇒>90∘\Rightarrow>90^\circ; chord ≤\leq diameter; r<r< distance to external point; fractions sum to 1; magnitude plausible against options. Name the licensing relation per step. Then emit the letter.

The appendix preserves the model-generated Primers without factual correction; they may contain inaccurate or incompletely qualified statements.

BibTeX

@misc{lu2026universal,
  title  = {Universal Textual Teaching for LLMs},
  author = {Lu, Zhanyi and Wang, Huan},
  year   = {2026},
  note   = {Manuscript}
}

Manuscript citation. The arXiv record will be linked when available.