TCM AI Four Examinations
Bring looking at the tongue and face, listening to the voice, and taking the pulse into one synthesized diagnosis — and make both the models and the conclusion acceptable






Before capture, lighting, framing, and occlusion rules are made clear so the model gets a usable image, rather than failing after the photo is taken.
After tongue capture, lighting, distance, and position all have to pass before analysis starts.
The top bar advances through face, tongue, listening, and pulse. Synthesis is an order inside one diagnosis, not four separate features.
Pulse uses a fingertip covering the camera. Finger method and force could not be controlled then, so the product states both the operation and the limit.
Examination analysis marks tongue color, teeth marks, lip color, and other abnormalities, and explains cold, heat, and blood stasis so the conclusion can be checked.Hover the preview and scroll to see content beyond the frame
The constitution report gives the conclusion first, then common signs and regulation advice, so someone without a TCM background can still see what to do next.Hover the preview and scroll to see content beyond the frame
Four examinations are not four photos
TCM diagnosis insists on four-examination synthesis: inspection (face and tongue), listening, inquiry, and pulse have to be read in the same diagnosis, not a conclusion from the tongue or a face alone. Many TCM AI products on the market can offer tongue and face as separate experiences, but cannot make one complete diagnosis.
These capabilities had to be handed to Ping An Health’s “Daily Suwen” and related businesses. The team delivered 25 models. My work was not claiming I trained them. It was making them capturable, checkable, and explainable, and together with the consultation system giving a conclusion users can understand.
Key product judgments
- Synthesis is a product principle, not a model slogan. The capture flow must bring inspection, listening, inquiry, and pulse into the same diagnosis.
- When the model is unsure, design a checking interaction first. Features that cannot be captured or captured well online are filled or checked with follow-up questions, instead of showing an uncertain recognition result directly.
- Accuracy cannot carry delivery on its own. Without shared annotation standards and bad-case handling, doctors and algorithm cannot agree on the same result.
- Users need to see “why this conclusion.” The report has to show the reasoning path and explain features that look like they conflict.
How synthesis enters one diagnosis
The earlier self-diagnosis was coarse: inspection, listening, and inquiry did not look like a TCM visit; the result page was too thin; disease and constitution prediction were not truly synthesized. The in-visit flow was closed into one chain:
Capture inspection and listening together. After taking face and tongue photos, guide the user to read a passage with five-tone cadence. Offline, while a doctor hears the chief complaint, they are already watching complexion and expression, and listening for whether the voice is abnormal.
Use questions to calibrate the model. Some inspection and listening features cannot be captured online, or recognition accuracy is low, so follow-up questions fill the gap. When the model gives a suspicious conclusion, rules also check it. For example, if tongue diagnosis judges putrid coating (a layer of dirt on the tongue that should scrape off), ask “scrape with your front teeth — does it come off?” Doctor experience becomes an executable matching rule here, not a review comment that stays in a meeting.
Bring pulse into the loop, and write the limit clearly. At the time, lighting a finger with the flash could estimate heart rate and HRV fairly well, but could not control finger method, force, or the mapping from sites to organs, so pulse conclusions were thin. Close the synthesis chain first, then iterate. Do not pretend online pulse already equals an offline visit.
Predict only after four-examination information is gathered. Constitution and possible disease should be based on all captured results from inspection, listening, inquiry, and pulse — not each model issuing its own conclusion.
Why the conclusion has to be explainable
Compared with Yilu, JD Health, WeDoctor, and Medlinker (whose “consultation” is closer to triage and appointment) as well as Jingmaibao, Xun’ai AI, and Quark AI, the judgment was: the product could show TCM features more completely, but after the visit it lacked treatment principles and methods and regulation advice, and the result page’s information structure was also thin. Xun’ai’s tongue, face, and inquiry could only be experienced separately and could not synthesize.
The new report page is organized around three questions:
- Disease care: what the disease might be, what the pattern is, and how to treat.
- Feature analysis: what was seen in face, tongue, listening, and pulse.
- Constitution regulation: constitution tendency and daily advice.
The report has to show reasoning from symptoms to location and pattern. When tongue or pulse does not match the typical signs of that pattern, “pattern transformation” is used to explain — for example, typical would be yellow coating and a stronger pulse, but the actual finding is a pale tongue and a weak pulse, which may be excess turning to deficiency, not a system error. Treatment principles and therapy advice align with the National Administration of Traditional Chinese Medicine’s clinical paths and diagnosis-and-treatment plans, so post-visit content does not become generated text that cannot be checked.
When is a model considered delivered
For four examinations and disease diagnosis to enter the business, annotation standards, data quality, bad cases, acceptance language, and integration all have to be handled together. Model accuracy is only one ring.
I took part in productizing tongue, face, listening, pulse, and disease-diagnosis models: aligning requirements with the business, annotation standards with doctors, and bad cases with algorithm. Disputed samples were deposited back into annotation rules, not recorded as a one-off recognition error. When data quality slowed progress, I led design of a doctor annotation tool, lowering the cost of doctors joining algorithm work, lifting annotation efficiency 30%+, and supporting on-time model delivery.
Results and limits
The annotation-tool efficiency lift can be mapped directly to product design I owned. 25 models, and face/tongue 20%+ above competitors, are team delivery results. The synthesized flow, follow-up checking, and report explanation correspond to product design and how it landed. The page does not write competitor accuracy as a personal KPI.
MY CONTRIBUTION
- Brought inspection, listening, inquiry, and pulse from scattered recognition into one synthesized flow: capture, follow-up checking, result display, and business acceptance designed together
- When a vision model was uncertain, used a product question to check (for example, if coating was judged putrid, ask whether it can be scraped off), instead of handing a model score to the user
- Took part in productizing tongue, face, listening, pulse, and disease-diagnosis models, and built annotation standards, bad-case handling, and a doctor annotation tool
PROJECT OUTCOME
- Took part in productizing 25 models; face and tongue accuracy 20%+ above competing products on the market
- The doctor annotation tool lifted annotation efficiency 30%+, supporting on-time model delivery
Scope note: Model count and competitor comparison are team or company results. Personal contribution focuses on the synthesized flow, checking interaction, result explanation, and the annotation tool.