Names: letter first, no specials, < 32 chars, case-sensitive. A(row,col) rows split by ; · A(1,:) row · A(1,1:2:5) cols 1,3,5 · .* ./ .^ element-wise, * matrix (inner dims agree) · zeros ones eye rand · for i=1:n … end, if … elseif … else … end, odd test fix(x/2)~=x/2 · index starts at 1 → a(n+1).
Course matrix A = [12 10 13 15 16; 22 45 65 1 0; 22 33 41 23 45; 21 30 12 6 2; 1 0 0 1 7]: sum(A) = [78 118 131 46 70], sum(sum(A)) = 443, diag = [12 45 41 6 7], trace 111. [3,5].*[4,8] = [12,40].
[0 2 10 20 3 15] with +1 if > 10 else −1 → [−1 1 9 21 2 16]. Plot: t=0:0.01:2*pi (0:1:2π gives only 7 points).
uint8 0–255 (uint16 65,535); toy 4×4 histogram [5 4 3 4]. medfilt2(X,[3 3]) odd window → middle of 9: {0,20,0,127,112,100,128,135,0} → 100; {0,5,8,60,99,99,109,125,155} → 99; ECG 1-D [1 3]. 3-D: V(:,:,k), isosurface(…,15), daspect([1 1 .4]), plot3, trisurf.
Coefficients highest power first: x³+4x²+9x+16 → [1 4 9 16]. roots ↔ poly; conv = multiply ([1 2 3 4]⋆[1 4 9 16] = [1 6 20 50 75 84 64]); deconv; polyder → [3 8 9]; polyint → [0.25 1.333 4.5 16 0].
$$\text{polyfit: }\min_p\sum_i(y_i-P(x_i))^2\ \Rightarrow\ (V^TV)p=V^Ty\quad(\text{41 pts}\to p\approx[0.965,\,0.140,\,4.969]);\qquad\text{spline passes exactly through the data}$$Solution exists iff rank(A) = rank([A b]) = r; unique if r = n (det ≠ 0); infinite if r < n (pinv, rref); A x = 0 non-trivial iff rank < n. Over-determined consistent → exact; inconsistent → least squares. Slide system A = [1 2 3; 4 5 6; 7 8 0], y = [366; 804; 351] → x = [25; 22; 99] (y/A is a bug). [3 −4; 6 −8] singular. Circuit A = [1 −1 1; −1 1 −1; 4 2 0; 0 2 5], b = [0;0;8;9] → i = [1, 2, 1] A. Chemical CO₂ + H₂O → O₂ + C₆H₁₂O₆: null space t·[6 6 6 1]. dot = projection, cross = moment.
Series: periodic only ("main limitation"); transform: periodic and aperiodic ("main advantage"). syms t; fourier(exp(-t^2)) = √π e^{−w²/4}; fourier(exp(-abs(t))) = 2/(1+w²).
"Backslash is Gaussian elimination (LU); the inverse is only for theory." "A solution exists when b lies in the column space: rank(A) = rank([A b])." "Fitting minimises squared error and need not touch the points; interpolation must." "A square wave has only odd harmonics falling as 1/k, so its edges need infinite bandwidth." "Sampling replicates the spectrum every fs; keep the copies apart with fs ≥ 2fm, then low-pass to recover."
MATLAB indices start at 1 (Hist(i+1), a(n+1)). uint8 saturates at 255: convert to double. medfilt2 needs an odd window. y/A ≠ A\y. det ≈ 0 means numerically singular. Variable names are case-sensitive (items vs Items bug). Fourier series only for periodic signals.
Init loss $=\log C$ ($\ln10=2.30$). Example (3.2, 5.1, −1.7): $p=(0.13,0.87,0.00)$, $L=2.04$.
Cat example $=2.9$; init $=C-1$. Zero once margins hold (stops learning); softmax never stops. $2W$ keeps $L=0$ → need $\lambda R(W)$.
$\beta_1$ 0.9, $\beta_2$ 0.999, start lr $10^{-3}$ or $5\cdot10^{-4}$. Numerical gradient check: $(1.25322-1.25347)/10^{-4}=-2.5$.
downstream = local × upstream. add distributes, mul swaps, max routes. $q=x+y,\ f=qz$: $\partial f/\partial x=z=-4$.
$$y=xW:\quad \frac{\partial L}{\partial x}=\frac{\partial L}{\partial y}W^T,\qquad \frac{\partial L}{\partial W}=x^T\frac{\partial L}{\partial y}$$Sigmoid: $\sigma'=\sigma(1-\sigma)=0.73\cdot0.27=0.20$; $dw=[-0.2,-0.39,0.2]$, $dx=[0.39,-0.59]$.
3×32×32, 10×(5×5), S1 P2 → 10×32×32, 760 params, 768,000 MACs. "Same" $P=(K-1)/2$. Receptive field $1+L(K-1)$. VGG: three 3×3 = one 7×7, $27C^2$ vs $49C^2$. ResNet $H(x)=F(x)+x$.
TP: class ok and IoU ≥ τ. Duplicate = FP. No TN. Course table: Cat 0.6667, Dog 0.5, Bicycle 0.3333 → mAP 0.5000. Bicycle IoU 0.74 fails τ 0.75.
$W$: $4h\times(h+d)$. $\partial c_t/\partial c_{t-1}=\mathrm{diag}(f)$ → "uninterrupted flow, like ResNet". Clip for exploding.
$\sqrt D$: $\mathrm{Var}(q\cdot k)=D$ → avoid softmax saturation. Permutation-equivariant → positional encoding. Mask future with $-\infty$. Block = MHSA → +res → LN → MLP(D→4D→D) → +res → LN; 6 matmuls; $O(N^2)$. ViT: $N=HW/P^2$ patches (224/16 → 196 tokens of 768), CLS token, low inductive bias → needs big data.
"Training minimises a loss by gradient descent; backprop is just the chain rule on the computational graph. CNNs share small filters across positions; LSTMs keep a cell state whose gradient passes through an element-wise gate; attention lets every token read every other token in one step at $O(N^2)$ cost." Overfitting = low train / high test error; underfitting = both high and "cannot be fixed by more epochs".
Error ≤ $s/2$, variance $s^2/12$. Weights per-channel, activations per-tensor, bias int32 with $s_b=s_as_x$. PTQ (observe min/max) vs QAT (fake-quant nodes). MCU: int32 accumulator, shift >> 7.
w = !(~X*X)*~X*y (~ transpose, ! Gauss–Jordan inverse, pivot < 1e-6 → abort). (1,2),(2,2.5),(3,3.5): $m=4.5/6=0.75$, $c=1.17$. Quadratic (1,2),(2,3),(3,5): $\beta=[2,-0.5,0.5]$. $\begin{bmatrix}2&1\\5&3\end{bmatrix}^{-1}=\begin{bmatrix}3&-1\\-5&2\end{bmatrix}$. "Linear in parameters, nonlinear in features."
1000 mAh, 50 mAh every 2 h → 1.67 d. 1200 mAh, 40 mAh, 5 d → every 4 h. Uno 500 mAh: awake 50 mA → 10 h; asleep 0.1 mA → 30 days; inference 45 mA·0.8 s = 0.01 mAh. $\lambda=1/12$: $S(2)$ 0.846, $S(8.3)$ 0.5, $S(24)$ 0.135; run if $S<0.3$ (t > 14.4 h). Adaptive interval $60TE/B$: 180 s @5000, 900 s @1000. Morning 7 min + night 22.5 min → 67 events → 2010 mAh.
Raspberry Pi: Interpreter → allocate_tensors → set_tensor → invoke → get_tensor; detection input [1,320,320,3] uint8, outputs boxes/classes/scores, thr 0.3, EfficientDet-Lite0. Arduino only if model < 20 KB. Hierarchy: Arduino wakes Pi at 10–30 cm.
| UART | I²C | SPI | |
|---|---|---|---|
| wires | 3 (Rx,Tx,GND) | 2 (SDA,SCL) | 4 + n (SCK,MOSI,MISO,SS) |
| clock | none (baud) | master | master |
| duplex | full | half | full |
| speed | 9600/115200 | 100k/400k/3.4M | ~10 MHz |
| notes | RS-232 ±3–25 V, MAX232 | 7-bit → 128 addr; START SDA↓ while SCL high; ACK SDA low; 4.7 kΩ pull-ups | SS low selects; byte per 8 clocks |
Ride OFF→ON 50%→OFF after 20 s. Fan Idle→30%→60%→100%. Traffic Green 120 → Yellow 30 → 3 blinks (6 states) → Red 120; pedestrian PB only in Green with > 30 s left. Python: Enum + loop; (value+1) % 4.
"Quantisation stores each weight as an integer plus a shared scale and zero-point; integer inference needs only an int32 accumulator and a rescale." "Idle current dominates the battery: sleep and wake on interrupts." "$S(t)$ is a survival probability, not a density." "Least squares has a closed form, $(X^TX)^{-1}X^Ty$, so a microcontroller can learn without gradient descent."
DNA = text in A,C,G,T; variant = changed letter (SNV) or small insert/delete (indel); monogenic = one gene (CF, sickle cell, PKU, MPS I, Duchenne, FMF, β-thal, Alport, Bardet–Biedl); motif = short meaningful pattern (splice site, TF binding site); reverse complement = reverse string, swap A↔T, C↔G (AACG → CGTT); ClinVar/ClinGen = curated databases; ACMG/AMP = 5-tier interpretation guideline; gnomAD = healthy-population variants.
Conv1D filters = motif detectors (local, position-invariant); max pooling = "motif present somewhere"; BiLSTM = context both upstream and downstream; combination claims local + long-range + contextual dependencies. Class imbalance: ×500 augmentation + class-inverse weights ($w_c\propto1/n_c$) + stratified batches; evaluate with weighted F1 and AUC-PR (PR more informative than ROC when positives are rare).
"Single-gene diseases are diagnosed by finding which DNA spelling change is harmful; yield is 25–50%. The authors cut a 101-letter window around known variants, place it in random background 500 times with augmentation, and train a 1-D CNN (motifs) plus a bidirectional LSTM (context) with class-weighted cross-entropy and Adam. It reaches 94.7% accuracy, F1 0.93, AUC-PR 0.98 on their synthetic test set; MPS I and PKU are weakest. The honest limit: synthetic negatives and no external validation, so clinical performance is unknown."
$\partial\mathrm{KL}/\partial z_s=(p_s-p_t)/T$. $T\to\infty$ uniform; $T\to0$ one-hot. Softmax over feature dims is a heuristic; AS-2 justifies it.
Tactile sensors = camera inside a soft gel; optics, gel, light differ per device → same material, different images → models do not transfer. Words ("rough, soft, slippery") are sensor-agnostic → use a frozen language model as teacher; distil its embedding into a ViT; afterwards train only a tiny head per dataset/sensor. Language only at training time.
AS-1: richer seq2seq teacher (BART) gives richer supervision. AS-2: feature-level KD ≫ cosine (direction only, flat near alignment) ≫ DKD (logits need a shared classifier; modalities differ). AS-3: moderate α, T; too supervised (α 0.8 → 80.68) or extreme T worse. AS-4: small batch regularises; lr 3e-5 unstable, 1e-5 slow. UMAP: tighter clusters. Grad-CAM: texture/edges, specular regions.
"Robots need touch; tactile cameras differ, so models fail across sensors. The authors use language as a sensor-agnostic teacher: a frozen BART embeds touch descriptions, a ViT is trained to match them with a temperature-softened KL loss (T 3.5) mixed 0.25/0.75 with cross-entropy, then frozen; only a tiny head is trained per task. On 39K relabelled DIGIT samples (32 classes) it reaches 95% at 100 shots, +13.3% average cross-sensor gain to GelSight, 98.8% on HCT, and beats vision-only using touch alone."
Layers: conv1D 32×3 ReLU ×2 → maxpool 2 → dropout 0.6 → flatten → dense 32 ReLU → dense 7 softmax. Accuracy/precision/recall/F1 by position: belt 80.9/81.8/79.1/80.4; wrist 92.4/92.6/92.1/92.4; upper arm 92.9/93.6/92.1/92.8; jeans pocket 98.2/98.5/98.1/98.3; on-device identical. TFLM rewrites 1-D conv as ExpandDims→2-D conv→Reshape, input (1,100,4); size halved (464 KB → ≈212 KB float32 weights), fits 1 MB flash.
$T_{ACC}$ 17.29 ms/sample, $T_{AI}$ 256 ms, $T_{Inference}=1729+256=1985$ ms (AI ≈ 13%). $E_{Ble}$ = total with BLE on − total with BLE off per cycle: I 6.12, II 5.81, III 6.07, IV 0.15 mJ/inference. Start-up (Table 6): reset 3.279 mA·0.509 s = 5.507 mJ; wake 1.003; accel config 6.556 mA·0.560 s = 12.115; BLE config 6.712 mA·0.174 s = 3.854 → configs > 70% of 22.5 mJ; model adds nothing at start-up.
Transmission is the most energy-hungry process of an IoT node. Putting inference on the device replaces 5000 bytes per window by one byte, but a connected BLE link still exchanges 11-byte keep-alive packets every connection interval, so fewer payload packets (III) save nothing; only switching the radio off between bursts (IV) cuts radio energy 40×. The AI cost (256 ms) is paid back by the longer connection interval it allows.
Accuracy unchanged on device (98.2%). Power 23.5 → 19.9 → 20.1 → 18.3 mW; over 1 h 84.5 J vs 66.0 J. Battery life ∝ 1/I: a 240 mAh cell lasts 33.9 h (I) vs 43.2 h (IV), +27%. Extra benefits: latency, less congestion, LoRa option, encryption; cost: raw data lost → hybrid mode (I to collect, IV to save).
"A wearable usually streams raw accelerometer data to a server. The authors put a 53k-parameter 1-D CNN on an Arduino Nano 33 BLE through TensorFlow Lite Micro; it recognises seven activities at 98% from the pocket position and runs in 256 ms. With a current analyser they compare streaming every sample with sending one label, two labels per byte, or a buffer with the radio off. Edge inference plus radio buffering uses about 21% less energy; sending fewer packets on a live link saves nothing because BLE keep-alives dominate. The saving really comes from radio scheduling that edge inference enables; the sink and server are not counted and the model is not quantised."