Skip to content

mlearning/tflite-micro: make the tflm tool run real models - #3804

Open
Abhishekmishra2808 wants to merge 1 commit into
apache:masterfrom
Abhishekmishra2808:mlearning/tflm-tool-inference
Open

Abhishekmishra2808 wants to merge 1 commit into
apache:masterfrom
Abhishekmishra2808:mlearning/tflm-tool-inference

Conversation

@Abhishekmishra2808

@Abhishekmishra2808 Abhishekmishra2808 commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Summary

The tflm tool could only allocate a model: it knew nine operators, ran Invoke() on uninitialized input and never printed the result. The 64 KB MicroProfiler also lived on the 4 KB task stack.

This PR registers the operators a model uses (about 90 builtins), verifies the file with the FlatBuffers verifier, and adds:

  • -I print operators, input/output tensors and arena usage
  • -d comma separated input values, quantized automatically
  • -x raw input tensor file
  • -n repeat inference and print the average Invoke() time

Outputs are printed dequantized, with argmax. Inputs are restored before every run because the memory planner may reuse their buffers. The resolver, profiler and interpreter are allocated on the heap.

Impact

  • tflm can now inspect and run .tflite models from NSH on sim:tflm, no board needed.
  • Fixes a stack overflow (profiler on the task stack) and wrong results on repeated runs.
  • Existing options (-i, -a, -E, -C, -o, -p) are unchanged. Only tflm_tool.cc changes; no Kconfig or build changes.

Testing

Host: Ubuntu 24.04 (WSL2) x86_64, gcc 13.3.0. Target: sim:tflm (Makefile build, no warnings). Models are from the downloaded TFLM tree, shared through hostfs.

nsh> mount -t hostfs -o fs=/tmp/t /t
nsh> tflm -I -i /t/hello_world_int8.tflite
model: 2704 bytes, schema v3, 1 subgraph(s)
description: MLIR Converted.
operators: 3
  FULLY_CONNECTED x3
0 (id=0): size=16, offset=16, first_used=0 last_used=1
1 (id=1): size=16, offset=0, first_used=1 last_used=2
2 (id=2): size=16, offset=16, first_used=2 last_used=3
3 (id=3): size=16, offset=0, first_used=3 last_used=3
 0: ................0000000000000000................................................ (1k)
 1: 11111111111111110000000000000000................................................ (1k)
 2: 11111111111111112222222222222222................................................ (1k)
 3: 33333333333333332222222222222222................................................ (1k)
input[0] "serving_default_dense_input:0" INT8 [1,1] 1 bytes scale=0.024480 zero_point=-128
output[0] "StatefulPartitionedCall:0" INT8 [1,1] 1 bytes scale=0.008291 zero_point=5
arena: 1120 of 8192 bytes used
nxai done!
nsh> tflm -i /t/hello_world_int8.tflite -d 1.5708
0 (id=0): size=16, offset=16, first_used=0 last_used=1
1 (id=1): size=16, offset=0, first_used=1 last_used=2
2 (id=2): size=16, offset=16, first_used=2 last_used=3
3 (id=3): size=16, offset=0, first_used=3 last_used=3
 0: ................0000000000000000................................................ (1k)
 1: 11111111111111110000000000000000................................................ (1k)
 2: 11111111111111112222222222222222................................................ (1k)
 3: 33333333333333332222222222222222................................................ (1k)
"Event","Tag","Ticks"
0,FULLY_CONNECTED,0
1,FULLY_CONNECTED,0
2,FULLY_CONNECTED,0
"Unique Tag","Total ticks across all events with that tag."
FULLY_CONNECTED, 0
"total number of ticks", 0
output[0] "StatefulPartitionedCall:0" INT8 [1,1] 1 bytes scale=0.008291 zero_point=5
  [0] 1.003206
inference: 1 run(s), total 650 us, average 650 us
nxai done!
nsh> tflm -i /t/speech.tflite -a 32768 -d 0 -n 20
0 (id=0): size=4000, offset=0, first_used=2 last_used=3
1 (id=1): size=1968, offset=0, first_used=0 last_used=1
2 (id=2): size=1968, offset=4000, first_used=1 last_used=2
3 (id=3): size=16, offset=4000, first_used=3 last_used=4
4 (id=4): size=16, offset=0, first_used=4 last_used=4
 0: 11111111111111111111111111...................................................... (2k)
 1: 11111111111111111111111111...........................222222222222222222222222222 (4k)
 2: 00000000000000000000000000000000000000000000000000000222222222222222222222222222 (6k)
 3: 00000000000000000000000000000000000000000000000000000........................... (4k)
 4: ................................................................................ (1k)
"Event","Tag","Ticks"
0,RESHAPE,0
1,DEPTHWISE_CONV_2D,1
2,FULLY_CONNECTED,0
3,SOFTMAX,0
"Unique Tag","Total ticks across all events with that tag."
RESHAPE, 0
DEPTHWISE_CONV_2D, 1
FULLY_CONNECTED, 0
SOFTMAX, 0
"total number of ticks", 1
output[0] "labels_softmax" INT8 [1,4] 4 bytes scale=0.003906 zero_point=-128
  [0] 0.250000
  [1] 0.250000
  [2] 0.250000
  [3] 0.250000
  argmax: 0 (0.250000)
inference: 20 run(s), total 250123 us, average 12506 us
nxai done!
nsh> tflm -i /t/audio..tflite -E
Unsupported custom operator: SignalWindow
nsh> poweroff

-d 1.5708 (pi/2) gives 1.003206, the sine of the input. The 20 micro_speech runs give the same output as a single run. The last command reports an unsupported custom operator by name.

The tflm tool could only allocate a model: it knew nine operators,
ran Invoke() on uninitialized input, and never printed the result.
The 64 KB MicroProfiler also lived on the 4 KB task stack.

Register the operators a model uses (about 90 builtins), verify the
file with the FlatBuffers verifier, and add:

  -I  print operators, input/output tensors and arena usage
  -d  comma separated input values, quantized automatically
  -x  raw input tensor file
  -n  repeat inference and print the average Invoke() time

Outputs are printed dequantized with argmax. Inputs are restored
before every run because the planner may reuse their buffers. The
resolver, profiler and interpreter are allocated on the heap.

Signed-off-by: Abhishek Mishra <mishra.abhishek2808@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant