Abstract:
Objective Mesoscopic laser-scanning microscopy enables high-resolution imaging over a millimeter-scale field of view. However, the optimal focal position often varies because of sample unevenness, section-thickness variation, field curvature, and residual system errors, making a single focal plane unable to maintain uniform sharpness. Photomultiplier-tube-based point scanning adds further challenges: weak fluorescence lowers the signal-to-noise ratio, while scan-sampling timing mismatch can cause horizontal discontinuities and structural misalignment. These effects distort focus-measure curves, producing fluctuations, false peaks, or peak shifts. To address these issues, a region-wise autofocus scheme was developed for a large-field-of-view multichannel laser-scanning microscope, integrating image standardization, adaptive region-of-interest selection, multi-scale focus evaluation, coarse-to-fine focal search, and electrically tunable liquid-lens focusing. The scheme aimed to improve focus discrimination, noise robustness, and full-field imaging consistency.
Methods The experimental platform consisted of a 488 nm continuous-wave laser, resonant and galvanometer scanners, an imaging objective, a photomultiplier tube, and an electrically tunable liquid lens. The field of view was divided into five regions. Liquid-lens voltage switching was synchronized with regional scanning so that each region could be imaged near its optimal focal plane.
The autofocus procedure included input standardization, regional focal optimization, and dynamically focused acquisition. Valid segments were extracted from raw photomultiplier-tube signals and reconstructed into two-dimensional regional images. Multiple initially defocused frames were averaged to suppress random noise. Horizontal discontinuities were detected from abrupt changes in lateral gray-level profiles. A horizontal offset was estimated during one-time calibration and reused throughout the voltage search.
An adaptive region of interest was generated for each region. Scharr operators calculated horizontal and vertical gradients, and local averaging of the gradient magnitude produced a gradient-energy heatmap. A quantile threshold retained the top 15% of high-texture pixels. Morphological processing removed isolated responses and filled small holes. Connected-component analysis, area screening, and mean-energy ranking were then performed. The bounding rectangle of the component with the highest mean energy was selected as the region of interest and remained fixed during focal searching.
An adaptive multi-scale gradient focus measure was constructed. Gaussian scale-space images were generated at several scales, and gradient energy was calculated at each scale. A low-quantile estimate represented the noise floor. Weak gradient responses below the corresponding threshold were treated as noise-related components. Structural-significance scores were calculated from structural and noise energies. Scores from different scales were weighted and fused, followed by nonlinear enhancement to improve peak distinguishability. A coarse voltage search located the peak neighborhood, and a fine search determined the optimal liquid-lens voltage for each region.
Performance was evaluated using the BBBC006 U2OS defocus sequence. The proposed focus measure was compared with Tenengrad, discrete cosine transform, Laplacian variance, difference of Gaussians, and tight-framelet-feature methods. A Gaussian-noise-degraded sequence was used to assess low-signal robustness. The adaptive window strategy was compared with center-window, golden-ratio, and saliency-guided methods. Peak width, clarity ratio, steepness, defocus-region fluctuation, computation time, and a composite index were calculated. Mouse kidney tissue sections were imaged to compare uniform focusing with region-wise focusing. Ten repeated autofocus trials were also conducted.
Results and Discussions On the original U2OS sequence, the proposed focus measure produced a peak width of 2, a clarity ratio of 12.59, a steepness of 0.24, a defocus-region fluctuation of 0.07, and a composite index of 2.88. Its composite performance exceeded that of the best comparison method by approximately 29.1%. The narrow peak and low off-focus fluctuation indicated stronger focal-position sensitivity and better suppression of unstable responses. The improvement resulted from combining fine-texture information, coarse structural information, and noise-floor estimation across multiple scales.
Under low-signal conditions, conventional measures showed larger fluctuations, weaker peak prominence, or local false peaks. The proposed measure retained a clear peak at the reference focal plane and achieved the highest composite index of 1.76. Its clarity ratio was approximately 56.3% higher than that of the tight-framelet-feature method and 44.1% higher than that of Laplacian variance. Its defocus-region fluctuation was approximately 6.7% lower than that of the tight-framelet-feature method and 22.2% lower than those of Laplacian variance and difference of Gaussians. Multi-scale structural fusion therefore reduced the influence of random gradient spikes while preserving persistent image structures.
The adaptive region-of-interest strategy achieved a composite index of 5.13, approximately 39.4% higher than that of the saliency-guided method. Its single-frame computation time was 0.80 ms, corresponding to a reduction of approximately 93.2% compared with large fixed-window strategies. The heatmap-based method concentrated evaluation on texture-rich regions and reduced background dilution. It improved focus-curve discrimination while decreasing unnecessary computation.
In the mouse kidney section experiment, region-wise focusing improved off-center sharpness by up to 30.7% compared with uniform focusing. Sharpness in regions R0, R1, R3, and R4 increased by approximately 30.7%, 25.7%, 7.6%, and 13.3%, respectively. The central region R2 changed by only about 1.6%. These results confirmed that a focal plane optimized for the center could not compensate for position-dependent focal differences across the full field. In ten repeated trials, the optimal voltage was (49.00 ± 1.56) V relative to a 50 V reference focal position. The relative standard deviation was approximately 3.2%, and the fluctuation relative to the reference position was approximately 3.1%, indicating good repeatability.
The complete calibration required several minutes because image acquisition and multi-frame averaging dominated the time cost. After regional optimal voltages were stored, dynamic focusing required only rapid voltage switching according to the scanning order.
Conclusions The region-wise autofocus scheme integrates multi-frame averaging, one-time horizontal misalignment compensation, adaptive heatmap-based region selection, multi-scale gradient evaluation, coarse-to-fine voltage search, and liquid-lens focusing. It addresses focal-plane inconsistency, low-signal reconstruction, and scan-sampling mismatch within one system framework. The method provides stronger focus discrimination, improved noise robustness, reduced window-based computation, better off-center sharpness, and stable repeated focusing. It is suitable for fixed tissue sections and other large-field-of-view photomultiplier-tube scanning tasks that permit regional focal calibration before formal acquisition.