Viva Prep - Shamma colour cheat sheet · 6 modules, blocks stacked one below the other · A4. Print with "Background graphics" enabled.
1

Analysis and Computing foundations

MATLAB for signals and CT images · polynomials · linear algebra · Fourier and sampling
formulanumbers to quotetrap / critiquesay this in the exam

MATLAB essentials

Names: letter first, no specials, < 32 chars, case-sensitive. A(row,col) rows split by ; · A(1,:) row · A(1,1:2:5) cols 1,3,5 · .* ./ .^ element-wise, * matrix (inner dims agree) · zeros ones eye rand · for i=1:n … end, if … elseif … else … end, odd test fix(x/2)~=x/2 · index starts at 1 → a(n+1).

Course matrix A = [12 10 13 15 16; 22 45 65 1 0; 22 33 41 23 45; 21 30 12 6 2; 1 0 0 1 7]: sum(A) = [78 118 131 46 70], sum(sum(A)) = 443, diag = [12 45 41 6 7], trace 111. [3,5].*[4,8] = [12,40].

Rectangle rule & loops

$$\int_a^b f\,dt\approx\sum_k f(t_k)\,\Delta t\qquad \int_0^\pi\sin t\,dt:\ \Delta t=0.5\to1.98,\ \Delta t=0.1\to1.9995,\ \text{exact }2$$

[0 2 10 20 3 15] with +1 if > 10 else −1 → [−1 1 9 21 2 16]. Plot: t=0:0.01:2*pi (0:1:2π gives only 7 points).

Images in MATLAB

sprintf names CT_001…012→imread → uint8 → double→hist_im: find(Y==i), Hist(i+1)→volume histogram (sum of slices)→valley T = 95→Z = Y if Y<95 else 255→imwrite(Z./255)

uint8 0–255 (uint16 65,535); toy 4×4 histogram [5 4 3 4]. medfilt2(X,[3 3]) odd window → middle of 9: {0,20,0,127,112,100,128,135,0} → 100; {0,5,8,60,99,99,109,125,155} → 99; ECG 1-D [1 3]. 3-D: V(:,:,k), isosurface(…,15), daspect([1 1 .4]), plot3, trisurf.

Polynomials & fitting

Coefficients highest power first: x³+4x²+9x+16 → [1 4 9 16]. roots ↔ poly; conv = multiply ([1 2 3 4]⋆[1 4 9 16] = [1 6 20 50 75 84 64]); deconv; polyder → [3 8 9]; polyint → [0.25 1.333 4.5 16 0].

$$\text{polyfit: }\min_p\sum_i(y_i-P(x_i))^2\ \Rightarrow\ (V^TV)p=V^Ty\quad(\text{41 pts}\to p\approx[0.965,\,0.140,\,4.969]);\qquad\text{spline passes exactly through the data}$$

Linear algebra

$$x=A\backslash b\ (\text{LU, not inv});\quad A^{-1}=\frac1{ad-bc}\begin{bmatrix}d&-b\\-c&a\end{bmatrix};\quad A^{-1}=\frac{\mathrm{adj}A}{|A|},\ C_{ij}=(-1)^{i+j}M_{ij};\quad x_{LS}=(A^TA)^{-1}A^Tb;\quad Av=\lambda v$$

Solution exists iff rank(A) = rank([A b]) = r; unique if r = n (det ≠ 0); infinite if r < n (pinv, rref); A x = 0 non-trivial iff rank < n. Over-determined consistent → exact; inconsistent → least squares. Slide system A = [1 2 3; 4 5 6; 7 8 0], y = [366; 804; 351] → x = [25; 22; 99] (y/A is a bug). [3 −4; 6 −8] singular. Circuit A = [1 −1 1; −1 1 −1; 4 2 0; 0 2 5], b = [0;0;8;9] → i = [1, 2, 1] A. Chemical CO₂ + H₂O → O₂ + C₆H₁₂O₆: null space t·[6 6 6 1]. dot = projection, cross = moment.

Fourier series, transform, modulation, sampling

$$x(t)=a_0+\sum_n\big[a_n\cos(2\pi nf_0t)+b_n\sin(2\pi nf_0t)\big],\ a_0=\tfrac1T\!\int_T\!x,\ a_n=\tfrac2T\!\int_T\!x\cos,\ b_n=\tfrac2T\!\int_T\!x\sin;\quad \pm1\text{ square: }a_0=0,\ b_n=\tfrac{4}{\pi n}\ (n\text{ odd})$$ $$X(f)=\!\int\! x(t)e^{-j2\pi ft}dt;\ \ \mathrm{rect}_\tau\leftrightarrow\tau\,\mathrm{sinc}(f\tau)\ (\text{zeros }n/\tau);\ \ e^{-at}u(t)\leftrightarrow\tfrac1{a+j2\pi f};\ \ e^{-a|t|}\leftrightarrow\tfrac{2a}{a^2+(2\pi f)^2};\ \ x\cos(2\pi f_0t)\leftrightarrow\tfrac12[X(f\!-\!f_0)+X(f\!+\!f_0)]$$ $$x_s=x\cdot\textstyle\sum_k\delta(t-kT_s)\ \leftrightarrow\ f_s\sum_kX(f-kf_s):\ \text{copies at }kf_s\pm f_m\ \Rightarrow\ \boxed{f_s\ge2f_m}\ \text{(200 Hz → 400 Hz; 150 → 300)};\ \text{RC: }\tfrac{V_o}{V_{in}}=\tfrac1{1+sRC}$$

Series: periodic only ("main limitation"); transform: periodic and aperiodic ("main advantage"). syms t; fourier(exp(-t^2)) = √π e^{−w²/4}; fourier(exp(-abs(t))) = 2/(1+w²).

Numbers to quote

sum(sum(A))443 · trace 111∫sin, Δt .5 / .11.98 / 1.9995 (exact 2)Lung thresholdT = 95 (valley of 12-slice histogram)Medians100 and 99polyfit[0.965 0.140 4.969] ≈ x²+5A\y[25 22 99]Circuiti = [1 2 1] ASquare waveb₁ 1.273, b₃ 0.424, b₅ 0.255; even 0Nyquistfs ≥ 2fmvar([2 4 6])4 (N−1 denominator)

Say this in the exam

"Backslash is Gaussian elimination (LU); the inverse is only for theory." "A solution exists when b lies in the column space: rank(A) = rank([A b])." "Fitting minimises squared error and need not touch the points; interpolation must." "A square wave has only odd harmonics falling as 1/k, so its edges need infinite bandwidth." "Sampling replicates the spectrum every fs; keep the copies apart with fs ≥ 2fm, then low-pass to recover."

Traps

MATLAB indices start at 1 (Hist(i+1), a(n+1)). uint8 saturates at 255: convert to double. medfilt2 needs an odd window. y/A ≠ A\y. det ≈ 0 means numerically singular. Variable names are case-sensitive (items vs Items bug). Fourier series only for periodic signals.

Viva cheat sheet · Analysis and Computingpage 1 of 6
2

Deep Learning foundations

MAI590/DEN790 · losses, optimisers, backprop, CNN arithmetic, detection metrics, LSTM, attention
formulanumbers to quotetrap / critiquesay this in the exam

Softmax & cross-entropy

$$p_k=\frac{e^{s_k}}{\sum_j e^{s_j}},\qquad L=-\log p_y,\qquad \frac{\partial L}{\partial s_j}=p_j-\mathbb 1[j=y]$$

Init loss $=\log C$ ($\ln10=2.30$). Example (3.2, 5.1, −1.7): $p=(0.13,0.87,0.00)$, $L=2.04$.

Multiclass SVM (hinge)

$$L_i=\sum_{j\neq y}\max(0,\,s_j-s_y+1)$$

Cat example $=2.9$; init $=C-1$. Zero once margins hold (stops learning); softmax never stops. $2W$ keeps $L=0$ → need $\lambda R(W)$.

Optimisers

$$\text{SGD: }w\leftarrow w-\alpha\nabla L\qquad\text{Momentum: }v\leftarrow\rho v+\nabla L,\ w\leftarrow w-\alpha v$$ $$\text{Adam: }m\leftarrow\beta_1m+(1-\beta_1)g,\ v\leftarrow\beta_2v+(1-\beta_2)g^2,\ \hat m=\tfrac{m}{1-\beta_1^t},\ \hat v=\tfrac{v}{1-\beta_2^t},\ w\leftarrow w-\alpha\tfrac{\hat m}{\sqrt{\hat v}+\epsilon}$$

$\beta_1$ 0.9, $\beta_2$ 0.999, start lr $10^{-3}$ or $5\cdot10^{-4}$. Numerical gradient check: $(1.25322-1.25347)/10^{-4}=-2.5$.

Backprop rules

downstream = local × upstream. add distributes, mul swaps, max routes. $q=x+y,\ f=qz$: $\partial f/\partial x=z=-4$.

$$y=xW:\quad \frac{\partial L}{\partial x}=\frac{\partial L}{\partial y}W^T,\qquad \frac{\partial L}{\partial W}=x^T\frac{\partial L}{\partial y}$$

Sigmoid: $\sigma'=\sigma(1-\sigma)=0.73\cdot0.27=0.20$; $dw=[-0.2,-0.39,0.2]$, $dx=[0.39,-0.59]$.

Convolution arithmetic

$$W'=\Big\lfloor\frac{W-K+2P}{S}\Big\rfloor+1,\qquad \#\text{params}=C_{out}(C_{in}K^2+1),\qquad \text{MACs}=C_{in}K^2\cdot W'H'C_{out}$$

3×32×32, 10×(5×5), S1 P2 → 10×32×32, 760 params, 768,000 MACs. "Same" $P=(K-1)/2$. Receptive field $1+L(K-1)$. VGG: three 3×3 = one 7×7, $27C^2$ vs $49C^2$. ResNet $H(x)=F(x)+x$.

Numbers to quote

ILSVRC top-5AlexNet 8L · VGG 7.3% · ResNet-152 3.57%CIFAR-1050k/10k, 32×32×3 = 3072Max-pool 2×2/2[[1,1,2,4],[5,6,7,8],[3,2,1,0],[1,2,3,4]] → [[6,8],[3,4]]Dropoutp 0.5; inverted: ÷(1−p) in trainingConv 7×7 backward∂L/∂b = 8; ∂L/∂W = [[20,10,2],[5,18,16],[15,10,4]]Transformer sizes12L/213M · GPT-2 48L/1.5B · GPT-3 96L/175B

Detection metrics

$$\mathrm{IoU}=\frac{|P\cap G|}{|P\cup G|},\quad P=\frac{TP}{TP+FP},\quad R=\frac{TP}{TP+FN},\quad AP=\sum_i\Delta R_iP_i,\quad \mathrm{mAP}=\tfrac1C\sum AP_c$$

TP: class ok and IoU ≥ τ. Duplicate = FP. No TN. Course table: Cat 0.6667, Dog 0.5, Bicycle 0.3333 → mAP 0.5000. Bicycle IoU 0.74 fails τ 0.75.

RNN → LSTM

$$h_t=\tanh(W_{hh}h_{t-1}+W_{xh}x_t)\quad\Rightarrow\quad \prod_t \mathrm{diag}(\tanh')W^T\ \text{vanishes/explodes}$$ $$\begin{pmatrix}i\\f\\o\\g\end{pmatrix}=\begin{pmatrix}\sigma\\\sigma\\\sigma\\\tanh\end{pmatrix}W\begin{pmatrix}h_{t-1}\\x_t\end{pmatrix},\ c_t=f\odot c_{t-1}+i\odot g,\ h_t=o\odot\tanh c_t$$

$W$: $4h\times(h+d)$. $\partial c_t/\partial c_{t-1}=\mathrm{diag}(f)$ → "uninterrupted flow, like ResNet". Clip for exploding.

Attention & ViT

$$Q=XW_Q,\ K=XW_K,\ V=XW_V,\qquad Y=\mathrm{softmax}\!\Big(\frac{QK^T}{\sqrt D}\Big)V$$

$\sqrt D$: $\mathrm{Var}(q\cdot k)=D$ → avoid softmax saturation. Permutation-equivariant → positional encoding. Mask future with $-\infty$. Block = MHSA → +res → LN → MLP(D→4D→D) → +res → LN; 6 matmuls; $O(N^2)$. ViT: $N=HW/P^2$ patches (224/16 → 196 tokens of 768), CLS token, low inductive bias → needs big data.

Say this in the exam

"Training minimises a loss by gradient descent; backprop is just the chain rule on the computational graph. CNNs share small filters across positions; LSTMs keep a cell state whose gradient passes through an element-wise gate; attention lets every token read every other token in one step at $O(N^2)$ cost." Overfitting = low train / high test error; underfitting = both high and "cannot be fixed by more epochs".

Viva cheat sheet · Deep Learningpage 2 of 6
3

Edge AI foundations

AIRE325 / MAI633 · quantisation, regression on MCUs, energy, TinyML, protocols, state charts

Quantisation

$$x_q=\mathrm{round}\!\Big(\frac{x}{s}\Big)+z,\qquad x\approx s(x_q-z),\qquad s=\frac{x_{max}-x_{min}}{255}$$ $$y_q=\frac{s_as_x}{s_y}(a_q-z_a)(x_q-z_x)+\frac{s_b}{s_y}b_q+z_y$$

Error ≤ $s/2$, variance $s^2/12$. Weights per-channel, activations per-tensor, bias int32 with $s_b=s_as_x$. PTQ (observe min/max) vs QAT (fake-quant nodes). MCU: int32 accumulator, shift >> 7.

Numbers to quote

Slide examplex 4.0 → x_q 102; a_q 64; b_q 13; y_q 112.8 → 11.05 (float 11.0)Exercises3.2/0.04 → 80 · 0.05(100−128) = −1.4 · 6.375/255 = 0.025Accuracy dropMobileNet 71.9→71.0 · KWS 94.0→93.8 · CIFAR 82.4→82.3Uno2 KB SRAM, 32 KB flash; int8+PROGMEM ≈ 4× neurons (50→200)Nano 33 BLE Sense256 KB / 1 MB; IMU, mic, T/H, pressureGesture lab119×6 samples, thr 2.5 g, 50-15-2, 600 ep, arena 8 KBKWS lab2500 ms @ 16 kHz, MFCC, Edge Impulse

Regression on a microcontroller

$$\hat\beta=(X^TX)^{-1}X^Ty\qquad m=\frac{n\sum xy-\sum x\sum y}{n\sum x^2-(\sum x)^2},\ c=\bar y-m\bar x$$

w = !(~X*X)*~X*y (~ transpose, ! Gauss–Jordan inverse, pivot < 1e-6 → abort). (1,2),(2,2.5),(3,3.5): $m=4.5/6=0.75$, $c=1.17$. Quadratic (1,2),(2,3),(3,5): $\beta=[2,-0.5,0.5]$. $\begin{bmatrix}2&1\\5&3\end{bmatrix}^{-1}=\begin{bmatrix}3&-1\\-5&2\end{bmatrix}$. "Linear in parameters, nonlinear in features."

Energy budgets

$$\text{uptime}=\frac{\text{capacity}}{\text{inf/day}\times\text{mAh/inf}},\qquad S(t)=e^{-\lambda t},\ \lambda=\tfrac1{\text{mean}},\qquad E[N]=\frac{T}{E[T]}$$

1000 mAh, 50 mAh every 2 h → 1.67 d. 1200 mAh, 40 mAh, 5 d → every 4 h. Uno 500 mAh: awake 50 mA → 10 h; asleep 0.1 mA → 30 days; inference 45 mA·0.8 s = 0.01 mAh. $\lambda=1/12$: $S(2)$ 0.846, $S(8.3)$ 0.5, $S(24)$ 0.135; run if $S<0.3$ (t > 14.4 h). Adaptive interval $60TE/B$: 180 s @5000, 900 s @1000. Morning 7 min + night 22.5 min → 67 events → 2010 mAh.

TinyML flow

train (PC)→prune · cluster · quantise→.tflite→xxd → model.h→TFLite Micro:ErrorReporterGetModel + versionOpResolver<N>tensor arenaInterpreterAllocateTensorsinput → Invoke → output

Raspberry Pi: Interpreter → allocate_tensors → set_tensor → invoke → get_tensor; detection input [1,320,320,3] uint8, outputs boxes/classes/scores, thr 0.3, EfficientDet-Lite0. Arduino only if model < 20 KB. Hierarchy: Arduino wakes Pi at 10–30 cm.

Protocols

UARTI²CSPI
wires3 (Rx,Tx,GND)2 (SDA,SCL)4 + n (SCK,MOSI,MISO,SS)
clocknone (baud)mastermaster
duplexfullhalffull
speed9600/115200100k/400k/3.4M~10 MHz
notesRS-232 ±3–25 V, MAX2327-bit → 128 addr; START SDA↓ while SCL high; ACK SDA low; 4.7 kΩ pull-upsSS low selects; byte per 8 clocks

State charts (method)

classify devices (in/out, dig/ana)→output string→count states→outputs/state→transitions (PB, After t, guard)

Ride OFF→ON 50%→OFF after 20 s. Fan Idle→30%→60%→100%. Traffic Green 120 → Yellow 30 → 3 blinks (6 states) → Red 120; pedestrian PB only in Green with > 30 s left. Python: Enum + loop; (value+1) % 4.

Say this in the exam

"Quantisation stores each weight as an integer plus a shared scale and zero-point; integer inference needs only an int32 accumulator and a rescale." "Idle current dominates the battery: sleep and wake on interrupts." "$S(t)$ is a survival probability, not a density." "Least squares has a closed form, $(X^TX)^{-1}X^Ty$, so a microcontroller can learn without gradient descent."

Viva cheat sheet · Edge AIpage 3 of 6
P1

CNN–BiLSTM for pathogenic variants

Abdelrehim & Mohamed 2026, Intelligent Systems with Applications 30:200654 · Liwa University

Pipeline

ClinVar / ClinGen variants→clean, title-case labels→101-bp window, random background ×500→reverse complement + k-mer jitter→+ random "Not a Disease"→A,C,G,T,N → 1,2,3,4,0 → one-hot→70/15/15 stratified, seed 42→multiscale Conv1D→max-pool→BiLSTM→attention/dense→softmax 21 classes

Equations

$$c_t^{(k)}=\sigma\Big(\sum_{i=0}^{f-1}w_{k,i}\cdot y_{t+i}+b_k\Big),\ t=1..L-f+1\ (1)\qquad p^{(k)}=\max_t c_t^{(k)}\ (2)$$ $$\overrightarrow d_t=\mathrm{LSTM}(\overrightarrow d_{t-1},y_t,\overrightarrow s_{t-1}),\ \overleftarrow d_t=\mathrm{LSTM}(\overleftarrow d_{t-1},y_t,\overleftarrow s_{t-1})\ (3,4)\qquad L=-\tfrac1N\sum_i\sum_c y_{i,c}\log\hat y_{i,c}\ (5)$$ $$\text{Spec}=\tfrac{TN}{TN+FP}\ (6)\quad \text{Acc}=\tfrac{TP+TN}{\text{all}}\ (7)\quad \text{Fallout}=\tfrac{FP}{FP+TN}\ (8)\quad LR^-=\tfrac{1-\text{Sens}}{\text{Spec}}\ (9)\quad NPV=\tfrac{TN}{TN+FN}\ (10)$$

Numbers to quote

Task20 monogenic diseases + "Not a Disease" = 21 classesWindow101 bp (±50), one-hot 101×4Augmentation×500 per mutation; RC + k-mer jitterTrainingAdam 1e-4 (β 0.9/0.999, ε 1e-8), batch 32, ≤50 ep, patience 7, best-val checkpoint, ≈12 ep, GPU, seed 42Resultsacc 94.7% · wF1 0.93±0.03 · AUC-PR 0.98±0.02 · ROC-AUC ≈1.00 · spec 0.98 · NPV >94%Hard classesMPS I (IDUA W402X), PKU (PAH R408W): acc 0.84, spec 0.889, 2 FP eachDiagnostic yield25–50% (exome 25–30%); ≈7,000 monogenic diseasesDeploymentFlask/Render web service "using LLMs"

Biology in one breath

DNA = text in A,C,G,T; variant = changed letter (SNV) or small insert/delete (indel); monogenic = one gene (CF, sickle cell, PKU, MPS I, Duchenne, FMF, β-thal, Alport, Bardet–Biedl); motif = short meaningful pattern (splice site, TF binding site); reverse complement = reverse string, swap A↔T, C↔G (AACG → CGTT); ClinVar/ClinGen = curated databases; ACMG/AMP = 5-tier interpretation guideline; gnomAD = healthy-population variants.

Why CNN + BiLSTM

Conv1D filters = motif detectors (local, position-invariant); max pooling = "motif present somewhere"; BiLSTM = context both upstream and downstream; combination claims local + long-range + contextual dependencies. Class imbalance: ×500 augmentation + class-inverse weights ($w_c\propto1/n_c$) + stratified batches; evaluate with weighted F1 and AUC-PR (PR more informative than ROC when positives are rare).

Critique (prepare calmly)

  • Negatives are random letters, positives are real motifs in random background → model may learn "real vs random". Authors admit it; fix: gnomAD/dbSNP benign variants.
  • 500 near-duplicate windows + random split → leakage between train/test; fix: split by mutation.
  • Ablation values printed as "[insert value]"; Table 2 rows have TP = 0 yet F1 0.93 (metrics from "hypothetical 19 negatives per class").
  • "Long-range" ≤ 50 bp; no benchmark vs CADD/SpliceAI; SHAP/PoSHAP and calibration are future work; LLM role unexplained.

Two-minute summary

"Single-gene diseases are diagnosed by finding which DNA spelling change is harmful; yield is 25–50%. The authors cut a 101-letter window around known variants, place it in random background 500 times with augmentation, and train a 1-D CNN (motifs) plus a bidirectional LSTM (context) with class-weighted cross-entropy and Adam. It reaches 94.7% accuracy, F1 0.93, AUC-PR 0.98 on their synthetic test set; MPS I and PKU are weakest. The honest limit: synthetic negatives and no external validation, so clinical performance is unknown."

Viva cheat sheet · Paper 1page 4 of 6
P2

Language-guided tactile representation learning

Mohsan, Ud Din, Xu, Abubakar, Hussain 2026 (arXiv 2609.14783) · KUCARS, Khalifa University

Pipeline

tactile image (DIGIT)→ViT student $f_\theta$ → $z_{tactile}$⇄ KLfrozen BART teacher → $z_{text}$+ CE on labels→$L=\alpha L_{CE}+(1-\alpha)L_{KD}$→freeze $\theta$→MLP head $h_\phi$ (0.017%)→material (touch only at test)

Equations

$$p_t^{(T)}=\mathrm{softmax}(z_{text}/T),\quad p_s^{(T)}=\mathrm{softmax}(z_{tactile}/T)\ (6)\qquad L_{KD\text{-}feat}=\mathrm{KL}(p_t^{(T)}\|p_s^{(T)})=\sum_j p_t\log\frac{p_t}{p_s}\ (7)$$ $$L_{student}=-\sum_c Y_T\log p_s(c)\ (8)\qquad L=\alpha L_{student}+(1-\alpha)L_{KD\text{-}feat}\ (9)\qquad \min_\phi \mathbb E[L_{CE}(h_\phi(f_\theta(x)),y)],\ \theta^*=\theta\ (10\text{–}12)$$

$\partial\mathrm{KL}/\partial z_s=(p_s-p_t)/T$. $T\to\infty$ uniform; $T\to0$ one-hot. Softmax over feature dims is a heuristic; AS-2 justifies it.

Numbers to quote

DatasetHCT 36K + SSVTP 4K = 39K DIGIT; 3 annotators; 32 classes: 12 distil / 20 unseenBest hyper-paramsα 0.25, T 3.5; fine-tune batch 32, lr 2e-5; RTX 4090, ≤100 epFew-shotK ∈ {0,10,100,1000}; ≈95% at 100-shotCross-sensorDIGIT → GelSight Hex: TAG +22% (0-shot), +18% (100), +9% (1000); ICRA18 +10.25%; avg +13.3%Table II (ours)TAG 61.20 · YCB 97.71 · ICRA18 73.86 · FEEL 86.88 · SSVTP 83.52 · HCT 98.80 · own 95.06Modalitieslang 60.38 · tactile 64.99 · vision 90.89 · ours 95.06AS-1 teacherBART 58.64 > DistilBERT 53.86 > RoBERTa 42.69AS-2 lossKD-feat 82.66 ≫ CosMin 58.64 ≫ DKD 34.02AS-4B 32 90.91 (512 → 45.44); lr 2e-5 95.06

Idea in one breath

Tactile sensors = camera inside a soft gel; optics, gel, light differ per device → same material, different images → models do not transfer. Words ("rough, soft, slippery") are sensor-agnostic → use a frozen language model as teacher; distil its embedding into a ViT; afterwards train only a tiny head per dataset/sensor. Language only at training time.

Ablation logic

AS-1: richer seq2seq teacher (BART) gives richer supervision. AS-2: feature-level KD ≫ cosine (direction only, flat near alignment) ≫ DKD (logits need a shared classifier; modalities differ). AS-3: moderate α, T; too supervised (α 0.8 → 80.68) or extreme T worse. AS-4: small batch regularises; lr 3e-5 unstable, 1e-5 slow. UMAP: tighter clusters. Grad-CAM: texture/edges, specular regions.

Critique (prepare calmly)

  • Distilled on one sensor; cross-sensor test classes overlap training classes (authors say so). Novel sensor + novel material untested.
  • Baselines train only on target data; proposed model gets 39K extra images → pretraining data vs language effect not separated (Table III is the cleanest evidence).
  • Zero-shot (K = 0) protocol with a trained head is unclear; ablation rows (58–83%) measured in a different setting than the 95.06% headline.
  • Language = short adjective lists (language-only 60%); may discard geometry/slip cues; vision-based sensors and recognition only.

Two-minute summary

"Robots need touch; tactile cameras differ, so models fail across sensors. The authors use language as a sensor-agnostic teacher: a frozen BART embeds touch descriptions, a ViT is trained to match them with a temperature-softened KL loss (T 3.5) mixed 0.25/0.75 with cross-entropy, then frozen; only a tiny head is trained per task. On 39K relabelled DIGIT samples (32 classes) it reaches 95% at 100 shots, +13.3% average cross-sensor gain to GelSight, 98.8% on HCT, and beats vision-only using touch alone."

Viva cheat sheet · Paper 2page 5 of 6
6

Paper 3 · Edge-AI cuts IoT power (activity recognition)

Muhoza, Bergeret, Brdys & Gary 2023, Internet of Things 24, 100930 · 1-D CNN on an Arduino Nano 33 BLE, four BLE scenarios

Pipeline

UTWENTE dataset: 10 people, 5 positions, 7 activities, 50 Hz→2 s windows, 50% overlap = 100 × (x, y, z, M)→DCNN 52,935 params, RMSprop, CE, batch 1024, 300 ep→TFLite Micro → C++ flatbuffer→Arduino Nano 33 BLE, 3.3 V, Keysight CX3324A→4 scenarios: energy vs streaming

Model & data equations

$$M=\sqrt{x^2+y^2+z^2};\quad f_s\ge2f_{max}\ (f_{max}<15\text{ Hz},\ 98\%\text{ power}<10\text{ Hz}\Rightarrow f_s>20\text{ Hz, used }50);\quad N_{win}=2\text{ s}\times50=100$$ $$\#\text{params}_{conv}=(kC_{in}+1)C_{out}:\ (3\cdot4+1)32=416,\ (3\cdot32+1)32=3104;\ 100\to98\to96\to48;\ 1536\cdot32+32=49184;\ 32\cdot7+7=231;\ \Sigma=52{,}935$$

Layers: conv1D 32×3 ReLU ×2 → maxpool 2 → dropout 0.6 → flatten → dense 32 ReLU → dense 7 softmax. Accuracy/precision/recall/F1 by position: belt 80.9/81.8/79.1/80.4; wrist 92.4/92.6/92.1/92.4; upper arm 92.9/93.6/92.1/92.8; jeans pocket 98.2/98.5/98.1/98.3; on-device identical. TFLM rewrites 1-D conv as ExpandDims→2-D conv→Reshape, input (1,100,4); size halved (464 KB → ≈212 KB float32 weights), fits 1 MB flash.

Energy model

$$E=V\,I\,\Delta t;\qquad E^{I}_{tot}=(E_{Acq}+E_{Ble,I})\tfrac{T}{T_{samples}};\qquad E^{(i)}_{tot}=(E_{Acq}+E_{AI})\tfrac{T}{T_{Inference}}+E_{Ble,(i)};\qquad E_{Acq}=V_{DD}I_{Acq}T_{Acq},\ E_{AI}=V_{DD}I_{AI}T_{AI}$$

$T_{ACC}$ 17.29 ms/sample, $T_{AI}$ 256 ms, $T_{Inference}=1729+256=1985$ ms (AI ≈ 13%). $E_{Ble}$ = total with BLE on − total with BLE off per cycle: I 6.12, II 5.81, III 6.07, IV 0.15 mJ/inference. Start-up (Table 6): reset 3.279 mA·0.509 s = 5.507 mJ; wake 1.003; accel config 6.556 mA·0.560 s = 12.115; BLE config 6.712 mA·0.174 s = 3.854 → configs > 70% of 22.5 mJ; model adds nothing at start-up.

Numbers to quote

Scenario Iraw 50 B/sample @ ~59 Hz, >28 kbps, CI 7.5 ms, 7.078 mA, 23.476 mWScenario II1-byte label / 2 s, <5 bps, CI 1500 ms, 6.035 mA, 19.915 mWScenario III2 × 3-bit labels / 4 s, CI 3500 ms, 6.007 mA, 20.054 mWScenario IVN-byte buffer, radio off, CI 7.5 ms burst, 5.554 mA, 18.329 mWSaving I→IV21.9% power ("up to 21%"); I→II 15%; II→III noneBLEadv 100 ms, CI 7.5–4000 ms, MTU 60 B, overhead 11 B, notify / write w/o responseDataset630,000 samples, ±2 g, 80/20 split, 20% of train for validationBoardCortex-M4, 1 MB flash, LSM9DS1, BLE 5 LE 2M PHYTable 8FPGA 97.5% @ 6.3 µW; KEH 79% saving; ours 98.3%, measured, edge vs cloud

Theory in one breath

Transmission is the most energy-hungry process of an IoT node. Putting inference on the device replaces 5000 bytes per window by one byte, but a connected BLE link still exchanges 11-byte keep-alive packets every connection interval, so fewer payload packets (III) save nothing; only switching the radio off between bursts (IV) cuts radio energy 40×. The AI cost (256 ms) is paid back by the longer connection interval it allows.

Results in one breath

Accuracy unchanged on device (98.2%). Power 23.5 → 19.9 → 20.1 → 18.3 mW; over 1 h 84.5 J vs 66.0 J. Battery life ∝ 1/I: a 240 mAh cell lasts 33.9 h (I) vs 43.2 h (IV), +27%. Extra benefits: latency, less congestion, LoRa option, encryption; cost: raw data lost → hybrid mode (I to collect, IV to save).

Critique (prepare calmly)

  • "Up to" 21% from one board, one dataset, one hour; no repeats or spread.
  • Sink (phone) and server energy excluded; no cloud variant with a smart radio schedule, so edge inference and radio management are conflated.
  • Float32 model (464 KB), no int8 quantisation; 59 Hz deployment vs 50 Hz training (100 samples = 1.73 s not 2 s).
  • Scenario IV latency grows with the buffer; apartment-bike data; young participants; phone sensors ≠ board IMU.
  • Strengths: precision measurements, transparent energy model, public dataset, lossless TFLM conversion, honest discussion.

Two-minute summary

"A wearable usually streams raw accelerometer data to a server. The authors put a 53k-parameter 1-D CNN on an Arduino Nano 33 BLE through TensorFlow Lite Micro; it recognises seven activities at 98% from the pocket position and runs in 256 ms. With a current analyser they compare streaming every sample with sending one label, two labels per byte, or a buffer with the radio off. Edge inference plus radio buffering uses about 21% less energy; sending fewer packets on a live link saves nothing because BLE keep-alives dominate. The saving really comes from radio scheduling that edge inference enables; the sink and server are not counted and the model is not quantised."

Viva cheat sheet · Paper 3page 6 of 6