Lab AI Assessment
Tender specs are only the entry ticket. Expand exam hardware into daily teaching and practice, then shift competition to scalable multimodal AI







School platform home, used by teachers to manage exams, lab classes, and practice resources.
The district teaching dashboard looks at this month’s classes, required-experiment coverage, and resource coverage — not only one exam.Hover the preview and scroll to see content beyond the frame
Regional exam score overview by subject, paper, and score band, for the education bureau’s post-exam analysis.Hover the preview and scroll to see content beyond the frame
Daily teaching: teachers turn a lab class into a flow that can be assigned, so devices enter ordinary lessons rather than exam week only. Older screenshot, used only to show the function.
Student practice: AI flags an operation error, gives the reason, a correct demonstration, and knowledge the student can keep asking about. Older screenshot, used only to show the function.
After practice and class, accuracy is broken down by action so evaluation can return to the next lesson, instead of stopping at a total-score report. Older screenshot, used only to show the function.
High-stakes exams still keep evidence and human review: the model detects, a person assigns the grade. Older screenshot, used only to show the function.
What this product does
The Ministry of Education required lab-operation exams to count in the senior-high entrance score. Cities then push physics, chemistry, and biology lab teaching and exam digitization so those exams can be traced and kept fair. Vendors help education bureaus write hardware and software specs, then win through public tender. So this is first a policy-driven, tender-delivered hardware-software system: city / district / school platforms, lab terminals, exam operations, AI scoring, human review, and device and compute operations sit on one chain.
It should not be a lab-recording tool, or special hardware that lights up a few times in the exam hall. Students operate at the bench; cameras have to see the whole process and technique; terminals have to be managed as classroom assets. AI has to face network, devices, compute, content versions, and scoring disputes in a real classroom.

I owned product planning and how the solution evolved. The competition was not only “are the features good.” It was three things that change by stage: whether parameters can enter the tender, why a school keeps using the system after buying it, and whether the solution can still scale when experiments and item types grow.
Three product judgments
-
Meeting tender specs is only the entry ticket. Policies differ by city, and so do the specs. Once several vendors can meet them, the real question is: why is our solution more worth buying? After devices enter the school, how do we raise actual use value, instead of serving a few exams a year?
-
Return on what was purchased matters more than one more exam feature. If the whole set is used two or three times a year, it is expensive idle capacity. Turn exam capability into infrastructure that daily teaching and student practice can also use. Frequency and teaching value go up, and tender competitiveness moves from “can examine” to “after the exam, it can still teach and still train.”
-
Once scenario coverage looks the same, the next variable is the underlying AI path. Teaching and practice gradually became table stakes. Stacking more scenes no longer creates a gap. Specialized CV models get expensive to label, train, and maintain as experiments and actions grow. Extensibility and delivery speed depend on whether a more general model can understand the experiment process.
Those judgments set the next three steps: first make exam hardware something schools will keep using; then, after feature competition flattens, change the technical base. I was not passively adding features as customers asked. At each competitive stage I kept asking what would actually move the solution next, and pushed the product there early.
How we broke through
01|Policy and tenders: match the specs, and explain how it will be used after purchase
The project is strongly policy-driven. When cities digitize lab teaching and exams, vendors help education bureaus write hardware and software specs, and procurement happens through tender. After comparing other vendors, my view was: the spec sheet decides whether you are eligible, not why a school chooses you.
So the solution cannot stop at “the exam can be traced and scored.” Tender materials, product planning, and on-site conversations all have to answer the same thing: after devices enter the school, besides those few exam days, how do they enter class, and how do they support student training.
The exam itself is still high-stakes. That is a floor, not a differentiator. Scoring must be reviewable, process must be traceable, exceptions must be recoverable — content snapshots, in-exam substitutions / make-ups, live monitoring, and human review exist so the exam can stand. Classroom terminals also have to be managed as device assets, or scaled delivery and on-site support costs will run away.
02|From exam hardware to daily teaching and student practice
Based on return after purchase, I pushed the team to fill two high-frequency scenes first. This was not adding two feature modules. It was changing how the whole solution gets used.
Daily teaching. Teachers can run a lab class on the devices, with AI as classroom support — scoring every student’s operations in real time, covering the pain that a teacher cannot coach every group. It also supports public classes and exemplary lessons powered by digitization and AI. A teacher usually only reaches a few groups in one period; real-time feedback fills in the rest of the class, while control of the lesson stays with the teacher. Teaching resources and exam resources are managed separately so practice state does not contaminate exam flow.

Student practice. For lab-operation training before the entrance exam, a self-study mode that can be repeated. Students complete operations on their own; AI recognizes the process in real time and gives evaluation and feedback. Devices can keep running outside exam windows and help improve real performance.
The product moved from “exam hardware used a few times a year” to teaching infrastructure that can keep entering the school’s daily flow. Exam, teaching, and practice states are mutually exclusive; the same terminals serve scenes with different reliability needs.
03|After scenes look the same, shift competition to the underlying AI
As the industry developed, teaching and practice modes became common vendor configuration. Scene coverage alone no longer created a clear gap. What would next affect product extensibility and delivery cost was the underlying AI path.
The original solution was mainly traditional CV deep-learning models. Different experiments and actions often needed their own data, labels, and training. As experiment count grew, or old experiments appeared as variants, model development and maintenance costs rose quickly, and new-project adaptation cycles stretched.
We then pushed the solution toward a small multimodal large model. Compared with the traditional path, a multimodal model depends less on new labels for new experiments or variant items; it does not need a separate model for every action; adaptation of new experiments and scoring rules can shorten. The AI capability moved from “train models for fixed item types” toward “understand the experiment process and evaluate it.”
This was not a one-time model swap. It was a redesign of how we would extend later, of algorithm cost, and of delivery speed on new projects. We started exploring a multimodal large model in place of many specialized CV models relatively early. Mainstream industry solutions later moved in a similar direction.
How the product is organized
For education bureaus, schools, teachers, students, and implementation ops, the product splits into city / district / school platforms, a teacher console, student terminals, grading review, and device and compute management.
The student terminal is the scoring site, not an ordinary tablet: capture conditions, identity checks, and device state are preconditions for scoring quality. The teacher console turns a lab period into an assignable flow. The review console shows process video, step rules, and a regrade entry together — the model detects, a person owns the grade. Class data is broken down by action so evaluation can return to the next lesson.
Exams, daily teaching, and student practice run on different flows and state constraints. Public classes were treated as product validation: missing capabilities went into the version, rather than remaining one-off support.
Results and limits
The product line supports regional teaching-and-exam programs and multi-school delivery, and in exam programs reached 98%+ human–AI agreement with a clear lift in grading efficiency. Those are whole-project results, completed together by product, algorithm, engineering, and delivery.
What I want seen is this: across policy tenders, school use value, and the underlying technical path, I kept judging which variable would actually move the solution next, and pushed the product there early.
MY CONTRIBUTION
- In policy-driven tender competition, judged that meeting specs is only the entry ticket; the real question is how schools keep getting value after they buy
- Pushed daily teaching and student practice into high-frequency use, turning exam hardware used a few times a year into teaching infrastructure that can enter the school’s daily flow
- After scenario coverage became similar across vendors, pushed a shift from many specialized CV models to a small multimodal model, cutting the cost of expanding, maintaining, and delivering new experiments and variants
- Built exam reliability into product mechanisms: content snapshots, in-exam exception handling, human review, and device-cluster operations
PROJECT OUTCOME
- The solution covers exams, daily teaching, and student practice, supporting regional teaching-and-exam programs and multi-school delivery
- In the Zhuhai senior-high entrance exam program, human–AI agreement reached 98%+, with a clear lift in grading efficiency
- Explored replacing many specialized CV models with a multimodal large model earlier than most peers; mainstream industry solutions later moved in a similar direction