← Back to experienceShipped product2022—2025Malong Technology

Lab AI Assessment

Tender specs are only the entry ticket. Expand exam hardware into daily teaching and practice, then shift competition to scalable multimodal AI

ROLEProduct planning and solution evolution
STAGELive and in scaled use
SURFACETenders & delivery · Scenario expansion
BOUNDARYOutcomes and personal contributions are shown separately

What this product does

The Ministry of Education required lab-operation exams to count in the senior-high entrance score. Cities then push physics, chemistry, and biology lab teaching and exam digitization so those exams can be traced and kept fair. Vendors help education bureaus write hardware and software specs, then win through public tender. So this is first a policy-driven, tender-delivered hardware-software system: city / district / school platforms, lab terminals, exam operations, AI scoring, human review, and device and compute operations sit on one chain.

It should not be a lab-recording tool, or special hardware that lights up a few times in the exam hall. Students operate at the bench; cameras have to see the whole process and technique; terminals have to be managed as classroom assets. AI has to face network, devices, compute, content versions, and scoring disputes in a real classroom.

Rows of student terminals and camera stands in a lab classroom, with students using Malong smart-lab terminals in the foreground
A real class: each group has a student terminal and a camera stand; the back wall shows class data. This is a hardware-software classroom, not tablets placed on a bench for the day.

I owned product planning and how the solution evolved. The competition was not only “are the features good.” It was three things that change by stage: whether parameters can enter the tender, why a school keeps using the system after buying it, and whether the solution can still scale when experiments and item types grow.

Three product judgments

  1. Meeting tender specs is only the entry ticket. Policies differ by city, and so do the specs. Once several vendors can meet them, the real question is: why is our solution more worth buying? After devices enter the school, how do we raise actual use value, instead of serving a few exams a year?

  2. Return on what was purchased matters more than one more exam feature. If the whole set is used two or three times a year, it is expensive idle capacity. Turn exam capability into infrastructure that daily teaching and student practice can also use. Frequency and teaching value go up, and tender competitiveness moves from “can examine” to “after the exam, it can still teach and still train.”

  3. Once scenario coverage looks the same, the next variable is the underlying AI path. Teaching and practice gradually became table stakes. Stacking more scenes no longer creates a gap. Specialized CV models get expensive to label, train, and maintain as experiments and actions grow. Extensibility and delivery speed depend on whether a more general model can understand the experiment process.

Those judgments set the next three steps: first make exam hardware something schools will keep using; then, after feature competition flattens, change the technical base. I was not passively adding features as customers asked. At each competitive stage I kept asking what would actually move the solution next, and pushed the product there early.

How we broke through

01|Policy and tenders: match the specs, and explain how it will be used after purchase

The project is strongly policy-driven. When cities digitize lab teaching and exams, vendors help education bureaus write hardware and software specs, and procurement happens through tender. After comparing other vendors, my view was: the spec sheet decides whether you are eligible, not why a school chooses you.

So the solution cannot stop at “the exam can be traced and scored.” Tender materials, product planning, and on-site conversations all have to answer the same thing: after devices enter the school, besides those few exam days, how do they enter class, and how do they support student training.

The exam itself is still high-stakes. That is a floor, not a differentiator. Scoring must be reviewable, process must be traceable, exceptions must be recoverable — content snapshots, in-exam substitutions / make-ups, live monitoring, and human review exist so the exam can stand. Classroom terminals also have to be managed as device assets, or scaled delivery and on-site support costs will run away.

02|From exam hardware to daily teaching and student practice

Based on return after purchase, I pushed the team to fill two high-frequency scenes first. This was not adding two feature modules. It was changing how the whole solution gets used.

Daily teaching. Teachers can run a lab class on the devices, with AI as classroom support — scoring every student’s operations in real time, covering the pain that a teacher cannot coach every group. It also supports public classes and exemplary lessons powered by digitization and AI. A teacher usually only reaches a few groups in one period; real-time feedback fills in the rest of the class, while control of the lesson stays with the teacher. Teaching resources and exam resources are managed separately so practice state does not contaminate exam flow.

A smart terminal on a lab bench scoring operations in real time, showing an overhead view and scoring steps
During class: the terminal films the bench and scores by action. The same hardware serves exams and ordinary lab lessons.

Student practice. For lab-operation training before the entrance exam, a self-study mode that can be repeated. Students complete operations on their own; AI recognizes the process in real time and gives evaluation and feedback. Devices can keep running outside exam windows and help improve real performance.

The product moved from “exam hardware used a few times a year” to teaching infrastructure that can keep entering the school’s daily flow. Exam, teaching, and practice states are mutually exclusive; the same terminals serve scenes with different reliability needs.

03|After scenes look the same, shift competition to the underlying AI

As the industry developed, teaching and practice modes became common vendor configuration. Scene coverage alone no longer created a clear gap. What would next affect product extensibility and delivery cost was the underlying AI path.

The original solution was mainly traditional CV deep-learning models. Different experiments and actions often needed their own data, labels, and training. As experiment count grew, or old experiments appeared as variants, model development and maintenance costs rose quickly, and new-project adaptation cycles stretched.

We then pushed the solution toward a small multimodal large model. Compared with the traditional path, a multimodal model depends less on new labels for new experiments or variant items; it does not need a separate model for every action; adaptation of new experiments and scoring rules can shorten. The AI capability moved from “train models for fixed item types” toward “understand the experiment process and evaluate it.”

This was not a one-time model swap. It was a redesign of how we would extend later, of algorithm cost, and of delivery speed on new projects. We started exploring a multimodal large model in place of many specialized CV models relatively early. Mainstream industry solutions later moved in a similar direction.

How the product is organized

For education bureaus, schools, teachers, students, and implementation ops, the product splits into city / district / school platforms, a teacher console, student terminals, grading review, and device and compute management.

The student terminal is the scoring site, not an ordinary tablet: capture conditions, identity checks, and device state are preconditions for scoring quality. The teacher console turns a lab period into an assignable flow. The review console shows process video, step rules, and a regrade entry together — the model detects, a person owns the grade. Class data is broken down by action so evaluation can return to the next lesson.

Exams, daily teaching, and student practice run on different flows and state constraints. Public classes were treated as product validation: missing capabilities went into the version, rather than remaining one-off support.

Results and limits

The product line supports regional teaching-and-exam programs and multi-school delivery, and in exam programs reached 98%+ human–AI agreement with a clear lift in grading efficiency. Those are whole-project results, completed together by product, algorithm, engineering, and delivery.

What I want seen is this: across policy tenders, school use value, and the underlying technical path, I kept judging which variable would actually move the solution next, and pushed the product there early.

MY CONTRIBUTION

  • In policy-driven tender competition, judged that meeting specs is only the entry ticket; the real question is how schools keep getting value after they buy
  • Pushed daily teaching and student practice into high-frequency use, turning exam hardware used a few times a year into teaching infrastructure that can enter the school’s daily flow
  • After scenario coverage became similar across vendors, pushed a shift from many specialized CV models to a small multimodal model, cutting the cost of expanding, maintaining, and delivering new experiments and variants
  • Built exam reliability into product mechanisms: content snapshots, in-exam exception handling, human review, and device-cluster operations

PROJECT OUTCOME

  • The solution covers exams, daily teaching, and student practice, supporting regional teaching-and-exam programs and multi-school delivery
  • In the Zhuhai senior-high entrance exam program, human–AI agreement reached 98%+, with a clear lift in grading efficiency
  • Explored replacing many specialized CV models with a multimodal large model earlier than most peers; mainstream industry solutions later moved in a similar direction