Can abstracted or encoded data be shared with AI assistants for preprocessing advice?

Hi organizers,

I’d like to clarify what information, if any, may be shared with external AI assistants such as ChatGPT or Claude when seeking help with data understanding and preprocessing.

I understand that competition data must not be sent to third-party AI services, and I have read the existing discussion on AI coding assistants and cloud compute. My question concerns locally derived representations rather than the original data files.

Would it be permissible to transform or summarize the training data locally into an abstracted or encoded representation, with the aim of preserving only the context needed to discuss preprocessing without exposing recognizable scans or patient-level records?

The purpose would be to better understand data characteristics, identify relevant methods or references, and get help designing or debugging preprocessing code. The original scans, image slices, identifiers, and patient-level labels would not be uploaded. This would be for development advice, not for processing test cases or generating submission predictions through an external AI service.

I do not assume that removing identifiers or encoding data automatically prevents reconstruction or makes sharing permissible. Could you clarify the boundaries for the following cases?

  1. High-level descriptions and aggregate statistics derived from the training data, without individual records.
  2. Sample-level transformed representations, such as extracted features or embeddings, where the original scan is not directly readable.
  3. Publicly available task information, independently created synthetic examples, and generic code containing no competition data.

Are any of these permitted, and what safeguards or conditions would apply? In particular, does the restriction cover all information derived from the competition data, or is there an acceptable level of abstraction for obtaining AI-assisted advice?

A few examples of what may and may not be shared would be very helpful. Thank you!

Hi @seonuk ,

Training data cannot be processed in any way that would permit reconstruction of the data or identification. From the cases you describe, case 1 (aggregate statistics over the whole training data set, and no information about individual images) is fine. Case 2 (sample-level transformed representations) would not be permitted because of the reconstruction and identification risk. Portions of case 3 (publicly available task information, generic code without any competition data or information about the data) would be acceptable, but independently created synthetic samples would not be, because the creation of synthetic versions of data based on individual samples may still carry the risk of reidentification. In general, questions that are general or aggregations over the whole set are ok, but anything on the individual record level should be avoided.