Open models, distilled to the size you ship.

INEWBIT TECH PTE LTD is a Singapore AI software and technology outsourcing firm. We distill open-source LLMs into smaller models, compress them for real hardware, and test and deploy them from cloud to edge.

Distillation flow An open-source source model flows to the right and becomes a smaller compact model, keeping its capability. open-source LLMs source smaller model compact capability kept Distillation flow An open-source source model flows to the right and becomes a smaller compact model, keeping its capability. open-source LLMs source smaller model compact capability kept

The work we take on

Three practice lines: distillation, compression, and edge deployment for open-source models.

Open-source LLM knowledge distillation

We transfer the capability of large open models into smaller models your team can own, fine-tune, and run. You receive the model weights, the distillation data pipeline, the training recipe, and the evaluation report.

  • Custom compact models trained from open-source models
  • Distillation datasets, synthetic data pipelines, and training recipes
  • Evaluation and regression reports with your quality bar
  • Full documentation so your team can reproduce every run
Llama Qwen DeepSeek LoRA / QLoRA SFT + alignment synthetic data
Knowledge distillation Capability moves from an open-source model through a funnel into a smaller model that keeps most of it. open-source LLMs source smaller model compact capability kept

Model compression and optimisation

We cut memory and latency while holding the quality bar you set. Quantisation, pruning, kernel tuning, and serving-side optimisations, validated against your benchmarks, not just ours.

  • Quantised and pruned checkpoints, ready for your runtime
  • Custom inference kernels for your stack and hardware
  • Serving configurations for vLLM, TensorRT-LLM, and more
  • Before-and-after benchmark reports, same inputs every time
GPTQ / AWQ INT8 / INT4 FP8 pruning vLLM TensorRT-LLM
Model compression A full-precision weight matrix is compressed into a lower-precision representation that uses far less memory. full precision heavier footprint quantised lighter lighter footprint

Edge AI testing and deployment services

We take models to the devices they will live on and prove they work there. On-device test matrices, benchmark harnesses, and release pipelines for phones, boards, and industrial hardware.

  • Device test matrices and benchmark harnesses per target
  • Runtime builds for NPU, CPU, and GPU targets
  • OTA-ready release pipelines with rollback plans
  • On-device performance and battery reports, per firmware
ONNX Runtime ExecuTorch RKNN Android / iOS embedded Linux
Edge deployment A compact model runs on the phone NPU on-device, with no cloud round-trip. model on-device measured on the device no cloud round-trip NPU runs on the silicon already in the device

How an engagement runs

Five steps, from first profile to a model your team owns and operates.

  1. 01

    Audit

    We profile your model, workload, hardware targets, and the quality bar that matters to your users.

  2. 02

    Distill

    We design the distillation setup and the data pipeline, then train and iterate with you.

  3. 03

    Compress

    Quantisation, pruning, and kernel work until the model fits your memory and latency budgets.

  4. 04

    Validate

    Benchmark and regression harnesses run on the exact hardware you ship, before you ship it.

  5. 05

    Deploy

    We hand over runnable artifacts, configs, and documentation, and support the rollout.

How we measure

No marketing numbers. Your quality bar becomes a test suite, and every deliverable has to pass it before you sign off.

Benchmark suites

Standard and custom evaluations, run with the same prompts and seeds before and after every change.

Regression harnesses

Every delivered checkpoint is tested against the acceptance suite agreed at the audit.

Device matrices

Edge work is measured on the exact devices and firmware your users run, not on reference rigs.

The company

INEWBIT TECH PTE LTD is a Singapore-registered company delivering AI software and technology outsourcing.

Our practice focuses on open-source LLM knowledge distillation, model compression and optimisation, and edge AI testing and deployment services.

  • Registered Singapore (Pte Ltd)
  • Business AI software and technology outsourcing
  • Practice Distillation, compression, edge testing and deployment