make: Entering directory '/mnt/fileserver/prj/dmctp/user/src/tensor/src/dev_tools/profiler' make: Nothing to be done for 'all'. make: Leaving directory '/mnt/fileserver/prj/dmctp/user/src/tensor/src/dev_tools/profiler' Tensor profiler: 25694 events, 16 graph nodes Summary capture span 12232.856 ms first-to-last kernel span 12154.662 ms kernel work (sum of kernel calls) 11771.017 ms library work (inclusive, overlaps) 12085.696 ms JIT compile work 235.731 ms Kernel execution operation dtype calls total_ms avg_ms MATRIX_INV_SQUARE REAL32 1 0.031 0.031 MATRIX_MUL REAL32 1 0.595 0.595 BLC UINT16 240 622.640 2.594 DPC UINT16 240 2258.614 9.411 AWB UINT16 240 790.590 3.294 DEMOSAIC UINT16 240 1367.619 5.698 CCM UINT16 240 1185.124 4.938 LUT UINT16 240 1122.410 4.677 BLUR UINT16 240 1577.172 6.572 BLEND_RGB2YUV UINT16 240 1807.102 7.530 CHROMA_SUBSAMPLE UINT16 240 1039.121 4.330 Library tasks (inclusive time) task dtype calls total_ms avg_us bytes mem_pool_init UNKNOWN 1 0.003 2.735 48 mem_alloc UNKNOWN 21 0.149 7.097 565069856 mem_view_copy UNKNOWN 4 0.018 4.571 262252 mem_pool_plan UNKNOWN 16 0.073 4.567 538549684 mem_pool_set_cleanup UNKNOWN 16 0.037 2.302 0 tensor_lazy_alloc REAL32 6 0.161 26.892 262324 mem_free UNKNOWN 32 0.070 2.174 565069856 tensor_lazy_op_dispatch REAL32 2 51.912 25956.022 0 tensor_alloc REAL32 6 0.057 9.445 262324 mem_copy UNKNOWN 4 0.147 36.788 262252 tensor_view_to REAL32 1 0.002 2.044 36 tensor_lazy_view_to REAL32 1 0.739 738.579 36 mem_view UNKNOWN 481 0.857 1.782 12630059520 tensor_lazy_alloc UINT16 10 0.209 20.874 538287360 tensor_lazy_op_dispatch UINT16 9 184.539 20504.360 0 tensor_alloc UINT16 10 0.135 13.468 538287360 tensor_view_to UINT16 480 0.966 2.012 28358553600 tensor_lazy_view_to UINT16 480 11819.333 24623.610 28358553600 tensor_free UINT16 10 0.052 5.194 538287360 tensor_free REAL32 6 0.031 5.245 262324 tensor_lazy_shutdown UNKNOWN 1 0.137 137.494 0 mem_pool_shutdown UNKNOWN 1 26.069 26069.297 577937408 JIT cache hits: 2403 cache misses: 11 hit rate: 99.5% task dtype calls total_ms avg_ms bitcode_bytes JIT PREPARE CACHE LOOKUP REAL32 2 0.010 0.005 0 tensor_op_00040002_8_8 REAL32 1 33.418 33.418 134884 tensor_op_00040001_8_8 REAL32 1 18.319 18.319 134884 JIT CACHE LOOKUP REAL32 2 0.009 0.004 0 JIT PREPARE CACHE LOOKUP UINT16 250 1.000 0.004 0 tensor_op_00010003_5_5 UINT16 1 16.895 16.895 134884 tensor_op_00010004_5_5 UINT16 1 25.249 25.249 134884 tensor_op_00010005_5_5 UINT16 1 17.067 17.067 134884 tensor_op_00010006_5_5 UINT16 1 23.272 23.272 134884 tensor_op_00010007_5_5 UINT16 1 19.379 19.379 134884 tensor_op_00010008_5_5 UINT16 1 16.265 16.265 134884 tensor_op_00010009_5_5 UINT16 1 24.497 24.497 134884 tensor_op_0001000a_5_5 UINT16 1 18.911 18.911 134884 tensor_op_0001000b_5_5 UINT16 1 22.459 22.459 134884 JIT CACHE LOOKUP UINT16 2160 8.803 0.004 0 Memory traffic and lazy pool external heap allocations 5 26520172 bytes external heap storage releases 5 26520172 bytes memory-block handle releases 32 565069856 logical bytes copy operations 8 524504 bytes moved pool block allocations 16 pool planned bytes 538549824 bytes pool committed bytes 577937408 bytes pool bump high-water 538549824 bytes pool reserved but unused 39387584 bytes committed-space utilization 93.2% plan realization 100.0% Operation graph node operation op_code dtype runs kernel_ms parents N0 ALLOC 0x00000000 REAL32 1 0.000 [] N1 ALLOC 0x00000000 REAL32 1 0.000 [] N2 MATRIX_INV_SQUARE 0x00040002 REAL32 1 0.031 [N0] N3 MATRIX_MUL 0x00040001 REAL32 1 0.595 [N1,N2] N4 ALLOC 0x00000000 REAL32 1 0.000 [] N5 ALLOC 0x00000000 UINT16 1 0.000 [] N6 BLC 0x00010003 UINT16 240 622.640 [N5] N7 DPC 0x00010004 UINT16 240 2258.614 [N6] N8 AWB 0x00010005 UINT16 240 790.590 [N7] N9 DEMOSAIC 0x00010006 UINT16 240 1367.619 [N8] N10 CCM 0x00010007 UINT16 240 1185.124 [N9,N4] N11 ALLOC 0x00000000 REAL32 1 0.000 [] N12 LUT 0x00010008 UINT16 240 1122.410 [N10,N11] N13 BLUR 0x00010009 UINT16 240 1577.172 [N12] N14 BLEND_RGB2YUV 0x0001000a UINT16 240 1807.102 [N12,N13] N15 CHROMA_SUBSAMPLE 0x0001000b UINT16 240 1039.121 [N14] kernel sequence (first execution of each node): N2:MATRIX_INV_SQUARE -> N3:MATRIX_MUL -> N6:BLC -> N7:DPC -> N8:AWB -> N9:DEMOSAIC -> N10:CCM -> N12:LUT -> N13:BLUR -> N14:BLEND_RGB2YUV -> N15:CHROMA_SUBSAMPLE Algorithmic throughput and tensor traffic kernel dtype class GOP/s tensor_GB/s OP/byte MATRIX_INV_SQUARE REAL32 floating-point 0.004 0.002 1.5000 MATRIX_MUL REAL32 floating-point 0.000 0.000 0.5000 BLC UINT16 integer 20.243 20.243 1.0000 DPC UINT16 integer 9.744 5.580 1.7461 AWB UINT16 integer 31.885 15.942 2.0000 DEMOSAIC UINT16 integer 36.863 18.432 2.0000 CCM UINT16 integer 55.834 31.905 1.7500 LUT UINT16 integer 75.797 33.744 2.2463 BLUR UINT16 integer 83.816 23.974 3.4961 BLEND_RGB2YUV UINT16 integer 66.259 31.386 2.1111 CHROMA_SUBSAMPLE UINT16 integer 7.581 27.291 0.2778 OP rates use the selected .op algorithm's static operations hint. tensor_GB/s is input+output tensor traffic, not hardware DRAM bandwidth. roofline data: tmp/isp_persistent_cpu_20260727.csv.roofline.csv roofline graph: tmp/isp_persistent_cpu_20260727.csv.roofline.svg kernel flame data: tmp/isp_persistent_cpu_20260727.csv.flame.folded