classify_card 2.0.1
classify_card: ^2.0.1 copied to clipboard
On-device Card Gate V3 pre-check: decides whether a photo holds an identity document before you call a classify/OCR backend. Runs offline in ~20ms, fails open, Android and iOS.
Changelog #
All notable changes to this project will be documented in this file.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
2.0.1 #
Replaces the classifier with the three-class Myanmar / Vietnam / other model, and rebuilds the scanner around it. 2.0.0 was never published, so these changes are folded in here rather than raising the major version again - but they are breaking, and a 2.0.0 that had shipped would have needed 3.0.0.
Breaking #
- The model is
myanmar-vietnam-other-foreign-v3-legacy-no-portrait-v1: EfficientNet-Lite0, INT8, three labels, one threshold per document class, temperature scaling instead of Platt. Policy schema 2. CardGateLabelhas three values -myanmarDocument,vietnamDocument,other- and carries the wire name the policy uses.isCardis deprecated in favour ofisAccepted.- Only a Myanmar document passes. A Vietnamese document is rejected exactly as
otheris, and the decision is made natively, where the label is produced, so nothing downstream can mistake one for a document it should accept.CardGateResult.scoresstill reports all three probabilities, so a caller can say "wrong country" rather than "no document". classifyandevaluatecrop the photo before inference. The model is trained on upright crops; handing it a whole photograph is the wrong input distribution, not a small loss of accuracy. Passcrop: falsewhen the image already is one.CardGateScannerViewwill not hand back a photo the classifier rejects, and checks when the countdown starts rather than when it finishes - about a second sooner, for the same single inference. A rejected document keeps the guide red while it is still in shot, measured by position rather than by a timer.- The viewfinder no longer classifies every frame. The corner detector drives the guide; the
classifier runs once per capture. Producing an upright crop per frame means a perspective warp
per frame - a full-resolution Bitmap on Android, a CIContext on iOS - and neither is affordable
at frame rate.
CardGateCapture.liveResultis therefore always null. - The scanner locks the device to portrait while it is up (
lockPortrait), and clears the preference on exit rather than guessing what the app had set.
Added #
CardGateFrameEncoder.clearCache()sweeps captured frames. They live in their own cache directory soDocumentCropper.clearCachecannot delete a photograph the host still holds a path to - which also means they needed their own sweep, and had none.DocumentCropper.cropacceptscorners, skipping the detector when the caller already knows where the document is. Worth about 860 ms on a mid-range phone.requireAcceptedDocument,verifyingHint,readyHint,startingHint,jpegQuality.
Notes #
- The bundled policy is an internal pilot:
pilot_only: true,production_ready=false, and its own release notes say not to use it to block a production request.CardGateMode.enforcetherefore reportsCardGateReason.pilotand allows the request through. That is enforced in code rather than documented and hoped for. - The Vietnamese threshold is not Wilson-certified in this release. It is moot while only Myanmar is accepted, but it is there in the policy.
2.0.0 #
Adds automatic document cropping, and turns the scanner into a hands-free capture screen. The
classifier is untouched: warmUp, classify, evaluate, classifyFrame and release keep
their exact signatures and behaviour, and their tests were not modified.
Breaking #
- Adds an ONNX Runtime native dependency, roughly 14 MB per ABI, and ONNX Runtime is arm-only:
DocumentCropperdoes not run on an x86 Android emulator. Everything else does. DocumentCropper.cropno longer takes anengine; there is one path and it has no UI.CardGateScannerViewhas no shutter and no Cancel button. The card is photographed once it has been held still inside the guide; the way out is a back arrow in the top-left corner, still wired toonCancel. Removed with them:autoCapture,autoCaptureFrames,shutterDisabledColor,cancelTextColorand theCardGateScanStatusenum.cardHintis nowholdingHint, joined byoutsideHint,tooFarHintandmovingHint.
Build floors are unchanged from 1.x: Android minSdkVersion 21, iOS 12.0.
Added #
DocumentCropper().crop(path)- a file goes in, a cropped file comes out, with no UI at all. Runs DocAligner'slcnet100_h_e_bifpn_256corner model (Apache-2.0) through ONNX Runtime. Measured against ground-truth corners on 24 real ID-card photographs plus 6 from a phone: median IoU 0.958, one photograph off by more than 30 px, and nothing cropped on the six containing no document. The geometric detector this package used before measures 0.872 with four such failures, and found nothing at all in any of the six phone photographs; it remains only as the fallback for when the model cannot load. Costs ~14 MB per ABI for ONNX Runtime plus 4.5 MB for the model, and ONNX Runtime is arm-only, so this path does not run on an x86 emulator. Every line of pre- and post-processing is Dart, once, so both platforms compute the same numbers by construction rather than by keeping two native ports in step.DocumentCropResultreports which detector ran, the four detected corners as fractions of the source photo, and why nothing was cropped when nothing was.
Added - hands-free capture #
CardGateScannerViewphotographs the card once the classifier sayscard, all four corners sit inside the guide filling at leastminFillRatioof it, and nothing has moved by more thanmaxDriftforholdDuration(default 1.2 s). A ring traces the guide as the countdown runs. Drift is measured against where the countdown started, not the previous frame: a slow creep passes every frame-to-frame check and still ends up somewhere else a second later.CardGateCornerSource- one method, four corners in upright frame pixels or null. The widget, the countdown and the geometry are detector-agnostic, so a different model drops in throughcornerSourcewithout touching any of them.- The photo is the live frame that completed the countdown, encoded natively - not a still
taken afterwards.
takePicturestops the stream, runs a fresh AF/AE convergence and commonly costs several hundred milliseconds on Android, and it photographs the scene at the moment the shutter fires rather than the one that was checked.resolutionPresettherefore sets the photo size as well, and now defaults toveryHighrather thanhigh;jpegQualityis new. CardGateFrameEncoder.encode(frame)exposes that path on its own.CardGateCapture.pathis the photo exactly as taken, uncropped. Cropping stays a separate API the caller invokes when it wants one:evaluatewas calibrated on whole photographs, and a capture screen that quietly did both would take that choice away.CardGateCapture.usedGeometryreports whether the corner detector was contributing. False means the hold ran on the classifier alone, which is what happens where ONNX Runtime cannot load.- A switch at the bottom of the scanner lets the user choose auto or manual. Manual gives them
a shutter and no countdown, and does not run the corner model at all - nothing is deciding when
to fire, so the guide only has to recolour, which is most of this screen's cost on a slow phone.
capture.modereports which took the photograph;showCaptureModeToggle: falselocks it toinitialCaptureMode. - The corner model runs on its own isolate, no more than once per
minCornerInterval(default 400 ms).OrtSession.runis a blocking FFI call: measured at ~690 ms per frame on a Galaxy A11 with another ~75 ms of post-processing, all of it on the isolate that draws the camera preview, which froze for three quarters of every second whenever a card was in shot. CardGateHoldTrackerandCardGatePreviewGeometryare separate, pure and tested on their own - 41 tests covering the countdown, the containment and fill rules, the drift rule and theBoxFit.covermapping between frame pixels and the guide, none of which need a camera to run.
Why this carries its own model #
ML Kit's entire public API is getStartScanIntent(Activity) -> Task<IntentSender>. No overload
accepts a Bitmap, Uri, File or InputImage, so it can only open its own camera UI - verified
three ways against the 16.0.0 AAR. Apple's VNDocumentCameraViewController is a view controller,
with the same consequence. Neither can be handed a file, so neither can serve crop(path).
Notes #
- Cropping never returns a worse photo than it was given. When no document is found,
result.pathis the path you passed in andresult.croppedis false. A wrong crop destroys information permanently; a missing crop only fails to add any, so every uncertain case resolves to leaving the photo alone.cropnever throws. - The detector was chosen by measuring four designs against the same 24 real ID-card photographs
and judging every crop by eye. The one shipped is the only one that produced no harmful crop:
21 of 24 correct, 3 declined. The three declined are clear cards it should have cropped - recall
is the known weakness, and it fails in the safe direction. The specification and the measurement
harness are in
tool/autocrop/. - Crops and scanned pages are written to cache directories owned by this package, never over your
input. The OS may clear them at any time, so upload or move the file rather than storing the path.
DocumentCropper.clearCache()andDocumentScanner.clearCache()delete only what this package wrote. - Cropping, scanning and classifying each have their own platform interface, their own error codes and their own thread. None can reach the others, so a host app can use any one of them alone.
1.0.2 #
Adds a realtime camera scanner. Everything that already worked is untouched: classify,
evaluate, warmUp and release keep their exact signatures and behaviour.
Added #
ClassifyCard.classifyFrame()scores a live camera frame with no disk write and no JPEG/PNG encode. Dart hands the plane bytes straight to the plugin; every colour conversion, rotation and resize happens natively on both platforms.CardGateScannerView, a camera screen with a card-shaped guide that recolours as the model scores frames, a shutter button, and anonCapturedcallback that hands back the photo path. What happens next - upload, OCR, your own API - is entirely the caller's business.CardGateOverlayShape, the guide itself, usable on its own.CardGateFrame.fromCameraImage()andCardGateException.codeBusy.
Notes #
- Frames are a viewfinder hint, not a verdict. The captured photo still goes through
classify/evaluate, which is the path the model was calibrated against. - The model runs on the whole frame, not the area inside the guide. Cropping to the guide would break the full-frame assumption the model was trained on; the guide is there to help the user aim.
- Backpressure drops frames rather than queueing them - queueing at frame rate grows without
bound and ends in an out-of-memory kill. A dropped frame comes back as
codeBusyand the scanner ignores it. - Adds a
cameradependency. If your app takescamerafrom a git fork, add adependency_overridesblock; see README.md.
1.0.1 #
- TODO: describe this release. No conventional-commit subjects were found since v1.0.0.
1.0.0 #
First public release of the on-device Card Gate V3 pre-check.
Added #
ClassifyCard.evaluate()returns a fail-openCardGateDecision. Only a confidentotherblocks a request; a missing plugin, a corrupt model, a decode error or a timeout all let the request through. Defaults toCardGateMode.shadowso a fresh integration cannot reject anyone.ClassifyCard.classify()for callers that want to own the error handling,warmUp()to avoid charging the first photo the cold-start cost, andrelease()to free the interpreter.CardGateResultexposeslabel, calibratedscore, the model's ownrawScore,modelVersion,latencyMsanddecodeMs.toMap()is safe to log - it never carries image data.- Model, policy and labels ship inside the package. The model's SHA-256 is verified against the policy on every startup, so a truncated or tampered copy makes the classifier unavailable rather than silently wrong.
- Example app with an
off/shadow/enforceselector for measuring the model against your own backend before enforcing.
Platform notes #
- Android runs
org.tensorflow:tensorflow-lite-task-vision:0.4.4, matching the vendor reference implementation exactly. - iOS runs
TensorFlowLiteSwift 2.10.0with equivalent preprocessing written by hand.TensorFlowLiteTaskVisioncannot be used in a host app: its static framework exports the sameTfLite*symbols asTensorFlowLiteC, it redefines MLKit'sGMLImage, and it has no arm64 simulator slice. See README.md. - Because the two platforms resample the photo through different code paths,
rawScorecan differ slightly between them for the same image. Measure on your own photos before switching toenforce.
Host app requirements #
These are one-line changes your app has to make:
- Android
minSdkVersion21 - TFLite Task Vision requires API 21, above the Flutter plugin template default of 19. - Android Gradle Plugin 7.4+ if your app also depends on anything pulling
androidx.annotation1.7.1 or newer; AGP 7.3 mis-resolves that Kotlin-multiplatform module. - iOS deployment target 12.0.
- iOS
Podfile:use_frameworks! :linkage => :static, becauseTensorFlowLiteSwiftships a static vendored framework. - Adds roughly 3.8 MB per platform to the host app.