# Optimizing Radiology Workflow: AI in Lung Nodule Classification and Breast Cancer Screening

## Learning Objectives

- Distinguish AI-based detection, classification, risk prediction, triage, and generative reporting functions.
- Interpret sensitivity, specificity, calibration, predictive value, and workflow outcomes in the context of screening prevalence.
- Apply AI output without departing from BI-RADS, Lung-RADS, Fleischner Society, and multidisciplinary management principles.
- Evaluate evidence from prospective breast-screening studies and externally validated pulmonary-nodule models.
- Identify automation bias, dataset shift, subgroup inequity, interoperability failure, and other deployment hazards.
- Design a monitored human–AI workflow with explicit thresholds, escalation pathways, and downtime procedures.
- Anticipate how longitudinal, multimodal, and opportunistic AI may alter preventive radiology.

---

## Introduction to AI in Radiology

<img src="images/fig_01.png" alt="Diagram of AI integration in radiology workflow">

Artificial intelligence in radiology is not one intervention. It is a family of narrowly defined tools that may detect a lesion, segment its boundaries, classify its morphology, estimate future disease risk, prioritize a worklist, retrieve prior examinations, draft report language, or identify a follow-up recommendation that has not been completed. Each task has a different reference standard, operating threshold, failure mode, and clinical consequence. A mammography model trained to mark suspicious regions cannot automatically be used to remove examinations from human review; a patient-level lung-cancer risk score cannot be treated as the malignancy probability of every pulmonary nodule.

Most contemporary image models use convolutional neural networks, vision transformers, or combinations of both. The model receives pixel or voxel data—often after resampling, intensity normalization, breast or lung segmentation, and series selection—and learns associations between imaging patterns and labels such as pathology, cancer registry outcomes, expert annotations, or longitudinal stability. Detection models return coordinates or heat maps; segmentation models return contours and volumes; classifiers return probabilities or categories. Longitudinal systems may register current and prior examinations to quantify growth or evolving asymmetry. Generative systems instead predict language or structured fields and therefore introduce distinct risks, including fabricated comparisons, omitted findings, and incorrect laterality.

**Teaching Point:** Performance is inseparable from the operating threshold. Lowering a threshold generally increases sensitivity while generating more false positives; raising it may reduce workload while missing disease. The clinically appropriate threshold depends on prevalence, the harm of delay, downstream testing capacity, and whether AI is being used as a safety net, concurrent reader, independent second reader, or autonomous triage gate.

A typical deployment routes selected DICOM series from the modality or PACS to an inference server. The result returns as a structured object, overlay, secondary capture, or worklist priority and must be reconciled with the correct patient, examination, series, and laterality. The system also needs visible failure states. “No result returned” must not be displayed or interpreted as “AI negative.” Latency, missing priors, incorrect series selection, unsupported implants, motion, and software-version mismatches can be as consequential as errors in the neural network itself.

**Framework:** Before acting on any output, define six elements: the intended task, applicable population, required input, validated threshold, permitted clinical action, and known failure modes. Then ask whether the model was evaluated on patients resembling the person in front of you—not merely on images from the same modality.

AUC and test-set accuracy do not establish clinical benefit. Useful evidence progresses from internal validation to geographically and technically independent validation, reader studies, silent deployment, prospective workflow trials, and patient-level outcome studies. Ayasa and colleagues’ review of AI across lung imaging, pathology, biomarkers, and precision oncology highlights both promising performance and persistent problems with heterogeneous datasets, limited external validation, standardization, and accountability ([PMID: 40943394](https://pubmed.ncbi.nlm.nih.gov/40943394/)).

**Nuance:** The unit of evaluation matters. An examination-level breast score can be correct while localizing the wrong lesion. A pulmonary-nodule algorithm tested only on surgically resected nodules sees a much higher prevalence and narrower spectrum than a screening program. A model trained using future reports may also contain label leakage—information that would not exist at the moment of clinical prediction.

AI should therefore be evaluated as part of a human–machine system. The relevant question is not “Is the model better than a radiologist?” but “Does this specific configuration improve decisions, timeliness, workload, equity, and outcomes without creating disproportionate harm?”

**MUST ACT:** Critical or suspicious findings remain subject to established communication and follow-up pathways. Do not allow AI availability, absence, or disagreement to delay direct radiologist review, diagnostic imaging, biopsy referral, or communication of an urgent result.

**Decision Point:** Choose the integration role based on the actual bottleneck. Worklist triage targets delay; concurrent marks target perceptual misses; an independent second reader targets double-reading workload; follow-up software targets loss to surveillance. One product should not be assumed to solve all four.

**Audience Poll:** Where is the largest preventable failure in your service—lesion detection, characterization, report turnaround, or closure of recommended follow-up?

---

## Case Study: AI Triage in Breast Cancer Screening

<img src="images/fig_02.png" alt="AI model flowchart for breast cancer screening triage">

AI can enter mammography screening by prioritizing worklists, identifying examinations likely to be normal, determining whether one or two human readers are required, marking suspicious regions, or supplying a safety-net alert when the radiologist and algorithm disagree. These functions are not interchangeable. A model validated for concurrent lesion detection should not automatically exclude examinations from review, and an examination-level malignancy score is not a BI-RADS assessment.

**Framework:** Treat AI as one input in a governed pathway: confirm technical adequacy; review current craniocaudal and mediolateral oblique views; compare with priors; incorporate symptoms, density, surgery, and risk history; assign the radiologist’s BI-RADS category; and route the patient accordingly. BI-RADS 1 or 2 returns to routine screening. BI-RADS 0 requires diagnostic evaluation. BI-RADS 3 generally follows a completed diagnostic workup and usually prompts short-interval surveillance. BI-RADS 4 or 5 requires tissue diagnosis. AI does not replace this accountability chain.

Consider an asymptomatic 60-year-old woman undergoing biennial screening. She has no previous breast cancer or known pathogenic germline variant, and her examination two years earlier was negative. Her breasts are heterogeneously dense. AI places the new examination in its highest-risk tier and localizes subtle architectural distortion in the posterior upper-outer left breast. The study moves forward in the worklist, but the radiologist first performs an independent review to reduce anchoring. Comparison with priors confirms that the distortion is developing.

The screening assessment is BI-RADS 0—not “cancer”—and she returns for diagnostic tomosynthesis with spot compression and targeted ultrasonography. The distortion persists and corresponds to an irregular 8-mm hypoechoic mass with angular margins. The completed diagnostic examination is BI-RADS 5. Ultrasound-guided 14-gauge core biopsy with clip placement demonstrates grade 2 invasive ductal carcinoma, estrogen-receptor positive, progesterone-receptor positive, and HER2 negative; axillary ultrasound is negative.

**Decision Point:** If the distortion disappeared on diagnostic views without a sonographic correlate, tissue summation would remain plausible, and escalation solely because of the AI score would be inappropriate. Conversely, persistent suspicious distortion without an ultrasound correlate should undergo tomosynthesis-guided biopsy rather than be dismissed because the ultrasound is negative.

**MUST ACT:** Never permit a high AI score to bypass diagnostic confirmation or a low score to override a suspicious human finding. A palpable mass, bloody discharge, new nipple retraction, suspicious calcifications, developing asymmetry, or architectural distortion requires conventional diagnostic evaluation regardless of algorithmic output.

Prospective evidence now supports carefully designed human–AI task sharing. In the randomized Swedish MASAI trial, AI triaged examinations to single or double reading and provided detection support. Among 53,043 women assigned to AI-supported screening, 338 cancers were detected, compared with 262 among 52,872 receiving standard double reading—6.4 versus 5.0 cancers per 1,000. Detection of invasive cancers increased from 217 to 270, predominantly through additional small, node-negative cancers. Recall was 2.1% versus 1.9%, without a significant increase in false positives, while screen-reading workload fell by approximately 44% (Hernström et al., *Lancet Digital Health* 2025; [PMID: 39904652](https://pubmed.ncbi.nlm.nih.gov/39904652/)).

PRAIM evaluated a different workflow in routine German double-reading practice. Across 12 sites, 119 radiologists screened 463,094 women; 260,739 examinations received AI-supported interpretation. Cancer detection was 6.7 versus 5.7 per 1,000, a 17.6% relative increase, while recall was slightly lower at 37.4 versus 38.3 per 1,000. The AI system could tag confidently normal examinations and issue a safety-net alert after a radiologist dismissed an examination considered highly suspicious (Eisemann et al., *Nature Medicine* 2025; [PMID: 39775040](https://pubmed.ncbi.nlm.nih.gov/39775040/)).

**Nuance:** MASAI was randomized; PRAIM was observational and allowed radiologists to choose the AI-enabled viewer, permitting selection and reader-behavior bias despite adjustment. Neither result automatically transfers to another vendor, population, density distribution, screening interval, or single-reader environment.

**Teaching Point:** Monitor the program, not algorithm sensitivity alone. Relevant measures include cancer detection, recall, false-positive and biopsy rates, interval cancers, stage and subtype distribution, reading time, and performance by age, density, race or ethnicity, scanner vendor, and screening round. Audit for automation bias, satisfaction of search after an AI mark, false reassurance from normal triage, and disproportionate detection of indolent ductal carcinoma in situ.

**Audience Poll:** Should your radiologists see AI localization before their initial review, only afterward, or through a documented two-pass reading protocol?

---

## Meta-Analysis of AI in Lung Nodule Classification

<img src="images/fig_03.png" alt="Statistical outcomes of AI lung nodule classification">

**Framework:** Separate three tasks that are frequently conflated. Detection localizes a candidate nodule. Classification estimates whether an identified nodule is benign or malignant. Risk prediction estimates whether the patient will develop lung cancer over a stated interval, sometimes from the entire CT rather than a segmented lesion. A patient-level risk score cannot automatically dictate management of a specific 7-mm nodule.

Wulaningsih and colleagues reviewed externally validated deep-learning computer-aided diagnosis models across 17 studies, 8,553 participants, and 9,884 CT-detected nodules. Pooled sensitivity was 0.88 and specificity 0.77. Deep-learning models were 11.6% more sensitive than physician judgment and 14.5% more sensitive than clinical risk models; specificity was similar to physicians and modestly higher than clinical models. Relative pooled AUCs were 1.03 versus physicians and 1.10 versus clinical models. Substantial heterogeneity means these estimates are not a guarantee of transportability ([PMID: 38782779](https://pubmed.ncbi.nlm.nih.gov/38782779/)).

**Teaching Point:** Benchmark the correct endpoint. Ardila’s end-to-end three-dimensional model achieved an AUC of 94.4% in 6,716 NLST cases and similar performance in an independent 1,139-case cohort. Without prior CT, it reduced false positives by 11% and false negatives by 5% compared with six radiologists; with priors, performance was comparable to the readers. This was scan-level cancer prediction, not proof that every highlighted nodule was correctly characterized ([PMID: 31110349](https://pubmed.ncbi.nlm.nih.gov/31110349/)).

Sybil predicts future lung-cancer risk from a single LDCT without clinical variables or radiologist annotations. One-year AUCs were 0.92 in held-out NLST data, 0.86 at Massachusetts General Hospital, and 0.94 at Chang Gung Memorial Hospital; six-year concordance indices ranged from 0.75 to 0.81 ([PMID: 36634294](https://pubmed.ncbi.nlm.nih.gov/36634294/)). These results establish discrimination, not calibration at every institution or improvement in patient outcomes.

**Nuance:** Reconstruct the differential before accepting a malignancy score. Spiculation, upper-lobe location, increasing solid component, emphysema, and growth increase concern for primary lung cancer. Alternatives include granulomatous disease, hamartoma, intrapulmonary lymph node, organizing pneumonia, scar, round atelectasis, septic embolus, carcinoid, and metastasis. Endemic histoplasmosis, tuberculosis, or coccidioidomycosis can sharply reduce positive predictive value without changing AUC. Benign calcification or fat, perifissural morphology, immune status, prior malignancy, multiplicity, and interval behavior can overturn the algorithm’s ranking.

**Decision Point:** Match the guideline to the acquisition context. Under [ACR Lung-RADS v2022](https://www.acr.org/-/media/ACR/Files/RADS/Lung-RADS/Lung-RADS-2022.pdf), a baseline solid nodule under 6 mm is category 2 with annual LDCT; 6 to under 8 mm is category 3 with six-month LDCT; 8 to under 15 mm is category 4A with three-month LDCT and possible PET/CT when the solid component is at least 8 mm; and 15 mm or larger is category 4B, generally prompting diagnostic evaluation. New nodules cross thresholds earlier: 4 to under 6 mm is category 3, 6 to under 8 mm category 4A, and 8 mm or larger category 4B.

For incidental nodules in adults at least 35 years old without known cancer or immunosuppression, Fleischner guidance uses clinical risk. A single solid nodule under 6 mm usually requires no follow-up, although 12-month CT may be considered in a high-risk patient. A 6–8-mm nodule warrants CT at 6–12 months and possible imaging at 18–24 months. A nodule over 8 mm warrants consideration of CT at approximately three months, PET/CT, or tissue sampling. Persistent pure ground-glass nodules at least 6 mm are followed for five years; part-solid nodules at least 6 mm require confirmation at 3–6 months and continued surveillance if the solid component remains under 6 mm (MacMahon et al.; [PMID: 28240562](https://pubmed.ncbi.nlm.nih.gov/28240562/)).

**MUST ACT:** Inspect thin-section images and priors. Automated volumetry is often more sensitive to growth than calipers, but approximately 25% volume change may be needed to exceed measurement variability. Slice thickness, reconstruction kernel, inspiration, motion, contrast, and vascular or pleural attachment can create false growth; subsolid and cavitary nodules are particularly segmentation-sensitive.

Nishida and colleagues studied 216 resected cN0 adenocarcinomas with pathological invasive diameter no greater than 30 mm—not an unselected nodule cohort. An AI-derived consolidation-volume/total-volume ratio of at least 0.72 predicted nodal metastasis. Among 117 tumors no larger than 20 mm, none below that threshold had nodal metastasis, and five-year recurrence-free survival was 100%. This finding is hypothesis-generating for resection planning, but its single-center retrospective design and surgical selection preclude using 0.72 as an autonomous operative rule ([PMID: 41212778](https://pubmed.ncbi.nlm.nih.gov/41212778/)).

**Audience Poll:** Which local failure is more dangerous: a missed aggressive cancer or an AI-driven cascade of CT, PET, biopsy, and surgery for benign disease?

---

## Assessing AI's Clinical Impact on Diagnostic Workflows

<img src="images/fig_04.png" alt="Flowchart of AI-driven clinical workflow optimization">

AI creates value only if technical performance changes care in a favorable direction. The causal chain begins with a valid input and calibrated output, proceeds through radiologist interpretation and downstream action, and ends with patient outcomes. Failure anywhere in that chain can erase an impressive test-set result.

**Framework:** Evaluate impact at four levels. Technical measures include inference failure, latency, segmentation quality, calibration, and performance across scanners. Diagnostic measures include sensitivity, specificity, false positives per examination, cancer-detection rate, recall, and interval cancer. Operational measures include reading time, turnaround time, backlog, arbitration, and follow-up completion. Patient measures include stage at diagnosis, unnecessary biopsy, complications, time to treatment, equity, and ultimately morbidity and mortality.

Predictive values are especially important in screening. In a hypothetical population with 1% disease prevalence, a model with 90% sensitivity and 90% specificity identifies nine true positives among 1,000 people but produces approximately 99 false positives; only about 8% of positive results are true disease. AUC hides this workload. Calibration asks whether patients assigned 20% risk actually have disease approximately 20% of the time. Decision-curve analysis goes further by testing whether acting at a threshold produces more net benefit than evaluating everyone or no one.

**Teaching Point:** Report both discrimination and consequences at the intended operating point. “AUC 0.94” does not tell the screening director how many women will be recalled, how many CT scans will be repeated, or how many benign nodules will be biopsied.

Workflow configuration determines impact. Worklist triage may shorten time to interpretation for suspicious studies but does not necessarily reduce total reading time. Concurrent marks may improve detection while increasing search interruptions. AI as a second reader can reduce double-reading workload but requires arbitration logic. Autonomous dismissal of low-risk examinations promises the greatest workload reduction and carries the greatest risk from false-negative triage. Follow-up tools may have modest diagnostic sophistication yet create substantial benefit by preventing surveillance loss.

**Decision Point:** Select the configuration that addresses a measured problem. If breast-screening sensitivity is acceptable but double-reading capacity is inadequate, second-reader substitution may be rational. If lung-nodule recommendations are accurate but frequently lost after discharge, closed-loop tracking may outperform another classifier.

Prospective evaluation should begin with a baseline period and silent deployment, during which outputs are logged but do not influence care. Investigators should prespecify thresholds, noninferiority or superiority margins, subgroup analyses, escalation rules, and stop criteria. A stepped-wedge rollout or randomized workflow trial is stronger than an uncontrolled before-and-after comparison because staffing, prevalence, scanner mix, and seasonal volume can change simultaneously. Follow-up must be long enough to capture interval cancers and benign resolution, not merely pathology from immediately biopsied lesions.

**Nuance:** Workflow metrics can improve while clinical quality worsens. A model may shorten median turnaround by moving high-scoring examinations forward yet create a long tail of delayed low-scoring cancers. Mean reading time may fall while arbitration and callbacks shift work to technologists or diagnostic clinics. Additional cancer detection may represent valuable stage shift, overdiagnosis, or both.

Human behavior must also be measured. Automation bias occurs when readers accept an incorrect suggestion; omission bias occurs when an unmarked lesion receives less scrutiny; alert fatigue reduces response to repeated false positives. Conversely, readers may distrust the model and duplicate every task, adding cost without benefit. Training should include local false-negative and false-positive examples and define whether the reader should perform an initial blinded pass before viewing AI.

**MUST ACT:** When AI and the radiologist disagree about a potentially consequential lesion, return to the images, priors, clinical context, and applicable guideline. Do not average the opinions. Determine which evidence explains the discordance and document the final human assessment.

A successful dashboard therefore tracks more than clicks and turnaround. It links AI output to radiologist decisions, diagnostic imaging, pathology, interval cancer, treatment, and follow-up. It also stratifies performance by clinically relevant subgroups. Only this closed chain can establish whether AI has optimized care rather than simply accelerated image processing.

**Audience Poll:** Which endpoint would persuade your service to continue an AI tool after one year—workload reduction, higher cancer detection, fewer interval cancers, faster diagnosis, or a favorable combination with no equity signal?

---

## Overcoming Challenges in AI Deployment for Routine Screening

<img src="images/fig_05.png" alt="List of common AI deployment barriers in radiology">

Routine screening is unforgiving: volumes are high, disease prevalence is low, and a small loss of specificity can generate thousands of recalls, examinations, or biopsies. AI should therefore be governed as a clinical intervention, not purchased as a software accessory. The intended-use statement must specify modality, population, acquisition protocol, task, user, output, and prohibited uses. “Detects nodules” is inadequate; “marks solid and subsolid nodules of defined size on adult noncontrast LDCT as a concurrent-reader aid” is testable.

**Framework:** Validate four linked layers: technical validity, reader performance, workflow performance, and patient outcomes. Excellent standalone discrimination does not establish that radiologists make better decisions, recalls remain acceptable, or patients reach diagnosis sooner.

Before activation, require independent external evidence and local silent validation. Local testing should represent actual scanner vendors, detector technologies, reconstruction kernels, dose levels, breast density or nodule morphology, prior surgery, implants, demographics, and screening prevalence. Report confidence intervals for sensitivity and specificity along with positive predictive value, false positives per examination, calibration, clinically consequential misses, and subgroup performance. Reader studies should compare unaided and AI-assisted interpretation and measure time and behavior, not only model-versus-radiologist accuracy.

**Decision Point:** Do not define success after examining the results. Prespecify acceptable margins for cancer sensitivity, recall burden, inference failure, and turnaround time, as well as who can pause the system when a boundary is crossed.

Monitoring must continue after launch. Dataset drift may follow scanner replacement, a new reconstruction algorithm, altered screening eligibility, or changing referral patterns. Dashboards should track input quality, missing outputs, score distributions, sensitivity among subsequently confirmed cancers, false-positive burden, calibration, and subgroup disparities. Delayed labels are essential because interval cancers and benign follow-up outcomes may take months or years to mature. Fairness audits should examine age, sex, race or ethnicity, breast density, smoking exposure, body habitus, disability-related positioning limitations, and access to prior imaging.

Interoperability failures are patient-safety failures. AI must receive the correct DICOM series, preserve patient and laterality identifiers, reconcile priors, return a reviewable result, and document model and version. Interfaces among modality, PACS, RIS, reporting software, electronic health record, and worklist orchestrator should fail visibly. Cybersecurity review should address encryption, least-privilege access, audit logging, vendor remote access, patch responsibility, software dependencies, data exfiltration, and corrupted inputs.

**MUST ACT:** Maintain a rehearsed downtime pathway. If the AI service or interface fails, examinations must automatically return to the standard human-read queue. No study should disappear, remain indefinitely “processing,” or be interpreted under the assumption that a negative AI result exists.

Human factors determine whether a technically sound system helps or harms. Poorly placed alerts cause fatigue; unexplained reprioritization may delay non-target emergencies; automation bias can cause readers to dismiss cancers that were not marked. Training should use local true-positive, false-positive, and false-negative cases and state explicitly that AI does not supersede BI-RADS, Lung-RADS, Fleischner guidance, or clinical judgment.

**Nuance:** Total cost includes integration, storage, cybersecurity, validation labor, monitoring, upgrades, incident response, downstream testing, and an eventual exit strategy—not only the license. Contracts should address data rights, uptime, model updates, change notification, indemnification, and access to performance logs. Regulatory authorization establishes a permitted intended use; it does not prove effectiveness in the local population.

Campanella and colleagues provide a useful deployment analogue from computational pathology. Their H&E foundation model for EGFR prediction achieved an AUC of 0.890 during prospective silent deployment, and modeled thresholds could reduce rapid molecular testing by as much as 43% while maintaining existing performance. The study was not a radiology-screening trial, but its staged validation, silent run, threshold selection, and resource-impact analysis are directly instructive ([PMID: 40634781](https://pubmed.ncbi.nlm.nih.gov/40634781/)).

A multidisciplinary oversight group should review incidents, drift, equity, updates, and continuing clinical benefit at a predefined cadence and after major changes. It must retain authority to restrict, roll back, or retire the model.

**Audience Poll:** Which failure would your institution detect first—falling sensitivity, rising recalls, subgroup disparity, or a broken interface?

---

## Future Directions for AI in Radiological Practice

<img src="images/fig_06.png" alt="Projection of AI future impact in radiological practices">

The next phase of radiological AI will move beyond isolated lesion detection toward longitudinal and multimodal decision support. Rather than returning “0.73 probability of cancer,” a clinically useful system may integrate current and prior imaging, report text, pathology, laboratory data, genomics, medications, and competing illness to recommend a bounded next step: retrieve an outside prior, perform targeted ultrasound, repeat CT in three months, discuss biopsy, or return to routine screening.

**Teaching Point:** Longitudinal reasoning may be more valuable than single-image accuracy. Subtle growth, increasing solid component, evolving architectural distortion, and imaging–pathology discordance often carry more information than morphology at one time point.

Foundation models trained across large image and text collections may support triage, segmentation, prior comparison, structured measurement, draft reporting, and cohort discovery with less task-specific labeling. Multimodal models could connect radiological phenotypes with histology or molecular testing. Evidence must nevertheless be described according to its true scope. Campanella’s EGFR model used digital H&E slides, not CT. Zhang and colleagues analyzed H&E whole-slide images from 517 patients with small-cell lung cancer, identifying histomorphological phenotypes and two prognostic subtypes; this was pathology-based risk stratification, not breast-screening triage ([PMID: 40898302](https://pubmed.ncbi.nlm.nih.gov/40898302/)). Such work shows how imaging-adjacent AI might route scarce confirmatory testing, but definitive molecular or pathological diagnosis still requires validated assays.

**Nuance:** Foundation models can fail broadly as well as generalize broadly. Fluent report language may contain fabricated prior comparisons, incorrect laterality, or omitted actionable findings. Generative output requires source grounding, uncertainty display, version control, and human verification.

Opportunistic screening is another high-value frontier. Existing CT examinations can quantify coronary calcium, emphysema, vertebral compression fractures, muscle mass, hepatic steatosis, or aortic size without additional radiation. Mammography may provide information about arterial calcification or future cancer risk beyond conventional density. The challenge is not measuring everything; it is selecting abnormalities with validated thresholds, an effective intervention, an accountable recipient, and closed-loop follow-up.

**Decision Point:** Add an opportunistic output only when four questions have answers: Is the measurement reliable? Does it change management? Who explains the result? How will completion of follow-up be verified?

Future lung programs may combine segmentation, volumetric doubling time, morphology, smoking exposure, prior cancer, circulating biomarkers, and comorbidity to personalize surveillance. Rare diseases and complex staging are also attractive targets. Benamore and colleagues’ review of malignant pleural mesothelioma describes imaging-based TNM updates and potential automated tumor-volume assessment, but not an autonomous staging solution; CT, PET/CT, MRI, pathology, and multidisciplinary review retain complementary roles ([PMID: 41339274](https://pubmed.ncbi.nlm.nih.gov/41339274/)).

AI may also improve operations by predicting incomplete examinations, protocoling studies, finding overdue surveillance, matching findings to guideline recommendations, and auditing whether action occurred. These lower-visibility applications may prevent more harm than another high-AUC detector. However, prediction of no-shows or nonadherence must not become a mechanism for deprioritizing disadvantaged patients.

**Framework:** The mature AI-enabled pathway should remain adaptive but bounded: validated inputs, calibrated risk, guideline-concordant options, explicit uncertainty, human authorization, closed-loop follow-up, and continuous outcome surveillance.

Research must progress from retrospective test sets to silent deployment, prospective reader studies, randomized workflow trials, and health-system outcome studies. Appropriate endpoints include interval-cancer rate, stage distribution, unnecessary biopsy, time to diagnosis, follow-up adherence, workload, equity, and cost. Mortality benefit may require prolonged follow-up, but intermediate outcomes should be prespecified rather than selected after results are known.

Privacy-preserving federated learning and distributed evaluation may expand representative datasets without centralizing every image, although neither eliminates label error, bias, or cybersecurity risk. Patients and institutions will also need transparent policies governing secondary findings, algorithmic prioritization, and reuse of images for model improvement.

**MUST ACT:** Build the outcome registry before scaling the algorithm. Without linkage to pathology, interval cancers, treatment, and follow-up, a department can measure speed but cannot determine whether patients benefited.

**Audience Poll:** Over the next five years, which capability would most improve your practice—longitudinal comparison, multimodal risk prediction, opportunistic screening, or automated follow-up closure?

---

## Case Scenario: Suspicious Lung Nodule Management

A 55-year-old current smoker with a 30-pack-year history undergoes baseline LDCT screening. He has mild emphysema but no fever, immune suppression, previous malignancy, or symptoms suggesting acute infection. Thin-section images reconstructed at 1 mm demonstrate a 9 × 8-mm solid right-upper-lobe nodule with irregular margins and focal spiculation. There is no macroscopic fat, benign calcification, or mediastinal adenopathy. AI segments the lesion at 310 mm³ and reports a 68% malignancy probability.

**Teaching Point:** Before considering the score, verify that the algorithm processed the intended thin-section series and that this patient falls within its validated population. Then independently inspect the entire CT. AI-directed attention to one nodule must not terminate the search for a second nodule, endobronchial lesion, pleural abnormality, or extrapulmonary finding.

Because this is a baseline solid nodule measuring 8 to under 15 mm, it meets Lung-RADS 4A criteria. Spiculation is an additional suspicious feature and may justify category 4X and more intensive diagnostic evaluation, but the AI probability itself does not change the Lung-RADS size definition. All available prior chest imaging should be retrieved; absence, stability, or previously unrecognized growth would materially alter risk.

The differential includes primary lung cancer, granuloma, focal scar, organizing pneumonia, and less likely metastasis or carcinoid. There are no infectious symptoms to support empiric antibiotics. Treating an asymptomatic indeterminate nodule with antibiotics can delay appropriate evaluation and provides little diagnostic information unless a genuine infectious syndrome is present.

**Decision Point:** Reasonable initial management includes three-month thin-section CT, with PET/CT considered because the solid component is at least 8 mm. PET sensitivity is limited for small lesions and indolent adenocarcinoma; a negative or mildly avid result cannot exclude malignancy. The Brock model can provide a transparent clinical estimate using age, sex, emphysema, nodule size and type, upper-lobe location, spiculation, family history, and nodule count. AI and Brock estimates should inform—but not replace—the guideline category and multidisciplinary judgment.

At three months, CT performed with comparable reconstruction parameters shows a volume of 430 mm³, a 39% increase. The approximate volume-doubling time is 190 days, supporting biologically meaningful growth after accounting for segmentation variability. FDG PET/CT shows only mild uptake and no nodal or distant disease. Pulmonary function testing demonstrates adequate operative reserve.

**MUST ACT:** Growth plus suspicious morphology requires escalation even if PET activity or a repeated AI score is low. Refer the patient for multidisciplinary review involving thoracic radiology, pulmonology, thoracic surgery, and, when relevant, oncology.

The team discusses navigational bronchoscopy, CT-guided biopsy, and surgical diagnosis. Choice depends on nodule location, expected diagnostic yield, pneumothorax risk, operative fitness, patient preference, and whether a nondiagnostic biopsy would change the plan. Tissue sampling confirms lung adenocarcinoma. Staging remains cT1aN0M0, and the patient undergoes an anatomic segmentectomy with systematic nodal evaluation. Pathology confirms a completely resected, node-negative stage IA tumor; adjuvant systemic therapy is not indicated, and CT surveillance is arranged.

**Nuance:** AI contributed by reproducibly segmenting the lesion and supporting risk stratification. It did not diagnose cancer, determine stage, select the biopsy route, establish operability, or choose the extent of resection. Those decisions required guideline context, longitudinal imaging, pathology, physiology, and patient preferences.

**Audience Poll:** At which point did this patient cross your threshold for tissue diagnosis—initial morphology, the AI probability, documented growth, or multidisciplinary synthesis?

---

## Tonight on Shift

- Confirm that the AI tool, patient population, modality, and processed series match the validated intended use.
- Review the complete examination and priors independently; never treat “no AI mark” as “no abnormality.”
- Translate output into BI-RADS, Lung-RADS, Fleischner, or other guideline-concordant management rather than acting on a raw score.
- Reconcile clinically important human–AI discordance using morphology, time course, patient risk, and diagnostic alternatives.
- Escalate suspicious findings through established communication, diagnostic imaging, biopsy, and referral pathways without AI-related delay.
- Close the loop: document the final assessment, ensure follow-up completion, and feed pathology, interval cancers, and false results into ongoing audit.
