-
Notifications
You must be signed in to change notification settings - Fork 31
All issues
Issue creation is restricted in this repository
Issues
is:issue state:open
is:issue state:open
Search results
[nightly-verify] main is red
priority:highHigh priorityHigh prioritystatus:readyReady to be worked onReady to be worked ontype:bugBug fixes, error corrections, or issue resolutionsBug fixes, error corrections, or issue resolutionsStatus: Open.#962 In lablup/mlxcel;fix(models): bound quantization params in the family-local MoE expert loaders
area:coremlxcel-core: MLX FFI, primitives, KV cache, layersmlxcel-core: MLX FFI, primitives, KV cache, layersarea:modelsModel architectures, weights, loading, metadataModel architectures, weights, loading, metadatapriority:mediumMedium priorityMedium prioritystatus:readyReady to be worked onReady to be worked ontype:bugBug fixes, error corrections, or issue resolutionsBug fixes, error corrections, or issue resolutionsStatus: Open.#958 In lablup/mlxcel;bug: Gemma 4 26B A4B standard-4-bit deterministically emits a wrong token (Middle War for Middle Ages) in long multi-turn context
priority:lowLow priorityLow prioritystatus:wontfixThis will not be worked onThis will not be worked ontype:bugBug fixes, error corrections, or issue resolutionsBug fixes, error corrections, or issue resolutionsStatus: Open.#933 In lablup/mlxcel;feat(xla): define operator-level MLX/IREE numeric contracts
area:architectureArchitecture and code structure changesArchitecture and code structure changesarea:coremlxcel-core: MLX FFI, primitives, KV cache, layersmlxcel-core: MLX FFI, primitives, KV cache, layersarea:inferenceGeneration, sampling, decoding (incl. speculative, DRY)Generation, sampling, decoding (incl. speculative, DRY)priority:highHigh priorityHigh prioritystatus:blockedBlocked by dependencies or other issuesBlocked by dependencies or other issuestype:enhancementNew features, capabilities, or significant additionsNew features, capabilities, or significant additionsStatus: Open.#932 In lablup/mlxcel;fix: KV estimation misses n_layer / n_embd for GPT-2 and GPT-BigCode
area:coremlxcel-core: MLX FFI, primitives, KV cache, layersmlxcel-core: MLX FFI, primitives, KV cache, layerspriority:lowLow priorityLow prioritystatus:doneCompletedCompletedtype:bugBug fixes, error corrections, or issue resolutionsBug fixes, error corrections, or issue resolutionsStatus: Open.#927 In lablup/mlxcel;fix(lora): make adapter weight fusion Conv1D-layout aware
area:inferenceGeneration, sampling, decoding (incl. speculative, DRY)Generation, sampling, decoding (incl. speculative, DRY)area:modelsModel architectures, weights, loading, metadataModel architectures, weights, loading, metadatapriority:lowLow priorityLow prioritystatus:doneCompletedCompletedtype:bugBug fixes, error corrections, or issue resolutionsBug fixes, error corrections, or issue resolutionsStatus: Open.#925 In lablup/mlxcel;Epic: Fused paged decode, sorting-free sampling, and serving performance techniques
area:architectureArchitecture and code structure changesArchitecture and code structure changespriority:highHigh priorityHigh prioritystatus:readyReady to be worked onReady to be worked ontype:enhancementNew features, capabilities, or significant additionsNew features, capabilities, or significant additionsStatus: Open.#909 In lablup/mlxcel;Mixed prefill/decode step execution: design spike and prototype
area:architectureArchitecture and code structure changesArchitecture and code structure changespriority:lowLow priorityLow prioritystatus:blockedBlocked by dependencies or other issuesBlocked by dependencies or other issuestype:enhancementNew features, capabilities, or significant additionsNew features, capabilities, or significant additionsStatus: Open.#908 In lablup/mlxcel;MLA matrix-absorbed decode path with compressed-latent KV cache for DeepSeek-family models
area:coremlxcel-core: MLX FFI, primitives, KV cache, layersmlxcel-core: MLX FFI, primitives, KV cache, layersarea:modelsModel architectures, weights, loading, metadataModel architectures, weights, loading, metadatapriority:lowLow priorityLow prioritystatus:readyReady to be worked onReady to be worked ontype:performancePerformance improvementsPerformance improvementsStatus: Open.#907 In lablup/mlxcel;Shape-bucketed kernel autotuner and cold-L2 benchmark methodology
area:benchmarkBenchmark harness and performance measurement (bench_*.sh, /update-benchmarks)Benchmark harness and performance measurement (bench_*.sh, /update-benchmarks)area:coremlxcel-core: MLX FFI, primitives, KV cache, layersmlxcel-core: MLX FFI, primitives, KV cache, layerspriority:mediumMedium priorityMedium prioritystatus:readyReady to be worked onReady to be worked ontype:enhancementNew features, capabilities, or significant additionsNew features, capabilities, or significant additionsStatus: Open.#906 In lablup/mlxcel;Fused residual-add RMSNorm and fused RoPE + KV-append decode kernels
area:coremlxcel-core: MLX FFI, primitives, KV cache, layersmlxcel-core: MLX FFI, primitives, KV cache, layerspriority:mediumMedium priorityMedium prioritystatus:readyReady to be worked onReady to be worked ontype:performancePerformance improvementsPerformance improvementsStatus: Open.#905 In lablup/mlxcel;Fused sparse-attention decode via page indirection for DSA and block-sparse models
area:coremlxcel-core: MLX FFI, primitives, KV cache, layersmlxcel-core: MLX FFI, primitives, KV cache, layersarea:modelsModel architectures, weights, loading, metadataModel architectures, weights, loading, metadatapriority:mediumMedium priorityMedium prioritystatus:blockedBlocked by dependencies or other issuesBlocked by dependencies or other issuestype:performancePerformance improvementsPerformance improvementsStatus: Open.#904 In lablup/mlxcel;