Intro to R Course
  • Prepare for the course
  • Copyright
  • Practical Sessions
  • Resources
  • Source Code
  • Report an issue
  1. Session 6 - Use of AI
  2. The coding assistant
  • Welcome
  • Session 1 - Basics of R
    • Getting familiar with RStudio
    • Setting up your Workspace
    • Functions that make the work
  • Session 2 - Tidyverse
    • Data manipulation using the Tidyverse
    • Logical conditions and Tidy
    • Creating variables
    • Grouping and summarising
  • Session 3 - Data Cleaning
    • Intro to Data Cleaning
    • Variable Class
    • Recoding variables
    • Derived Variables & Export
  • Session 4 - Tables
    • Counting cases
    • Crosstabulations and richer tables
    • Tables of things you cannot count
    • The whole table in one line
  • Session 5 - ggplot2
    • Scatterplot - your first plot
    • Barplots - elemental count
    • Lines - tracking trends
    • Histograms for Epicurves
  • Session 6 - Use of AI
    • The teaching assistant
    • The design assistant
    • The coding assistant

On this page

  • Part 6 · Understand and plan before prompting
    • Ask for a strategy, not a solution
    • Execute and verify - step by step
  • Part 7 · What you should have noticed
    • Capitalisation and whitespace
    • A symptom listed twice
    • Is rash the same as petechiae?
    • Two different kinds of “nothing”
  • Exercise summary
  1. Session 6 - Use of AI
  2. The coding assistant

The coding assistant

Session 6 practical exercises

An email at your inbox calls your attention. It comes from the Institute’s data science team, right to you:

“Sorry it took us more time than expected. Kassandra asked us to prepare more some data for you, now that we have our newest linkage to the clinical capture system to the surveillance registry in place. We have collected available clinical information about your cases, you’ll find it as a new column, reported_symptoms, in the dataset I’m sending you from the last cleaned version we received. The tool is still ongoing fine-tuning and bug-fixing, and since it’s based on free text entries from the clinicians, it may not be perfect tidy data yet. Good luck!”

Kassandra was cc’ed, and soon enough you receive a second mail:

“I want a symptom frequency count for the case definition review next week, send it to me before tomorrow so I have time to review it carefully during the weekend”

Doesn’t seem like a complicated petition, considering what you have already done so far… right?

Action — Open the new dataframe and explore the variable reported_symptoms. Can you produce that table right away?

Part 6 · Understand and plan before prompting

This is a different kind of mess to anything you have cleaned so far in this course. It is not a typo, not an inconsistent label, not a wrong data type. It is a structural problem: the information you need exists, but it is not organised the way you need it to answer Kassandra’s question.

You will use an AI assistant to help you solve this. But before you open that chat window, there is something worth sitting with for a moment:

  • What do you need to accomplish, exactly?

  • Which problems can you identify from the beginning?

  • How will you move from what you have to where you need to go?

The skill this exercise is really about is not any single R function, it is how you work with an AI assistant through a problem like this: how you frame it, how you check its answers, and how you come back with a sharper question when something breaks. A vague question to an AI assistant about data you do not understand tends to produce a vague, generic answer — one that solves a version of the problem, not necessarily yours.

So before writing a single prompt, spend a few minutes with the column on your own terms:

Action — Understand your starting point: explore reported_symptoms using tools you already know: unique(), count(), and manually inspecting a handful of raw rows. Take note of everything you notice: How many symptoms seem to appear per case? Is the order consistent? Are there any values that surprise you? This is your starting point.

Action — Define the destination before the path: before asking for any code solution, think about how your data should look like once cleaned.

With this in mind, we can go back to our AI tool of choice.

Ask for a strategy, not a solution

By combining the information from your starting point and your desired destination, the first prompt emerges organically: how do I arrive from A to B? This is a pathway, not a code solution. We will deliberately avoid asking for code now, that comes later.

Action — Start building your prompt, including:

  • Contextual baselien information: who you are, what are you doing, how do you want AI to reply, with which libraries or approaches, in which tone. You should ask explicitly AI to use Tidyverse, and explain you the proposed solutions (functions, logic, etc.) by the way.

  • Your specific problem: starting and destination points

  • Your petition: a step-by-step strategy, in plain language first without code

Action — Check the proposed strategy against what you noticed when exploring your data. Does it account for what you saw? Would it handle a case with only one symptom? A case with none?

If the first suggestion sounds too simple

A very common first suggestion is to split the string into a fixed number of columns — symptom_1, symptom_2, and so on. Think back to what you noticed about ordering before accepting that. Does the first value in the string always mean the same thing across cases?

Action — A little back-and-forth to refine the plan is usually needed. Check the plan, and prompt again if you are not convinced: identify the problematic bits and ask for correction, or explanations until the plan is set

The AI tool is only as good as your prompt, but it is also quite clever for spotting inconsistencies. Is the strategy it proposed seem truly odd and far from your original ideas, ask why, and you may be surprised.

Execute and verify - step by step

This is the part where the plan meets the actual data.

Action — Now ask AI for the code, once the plan is agreed and fixed.

Do not run the whole strategy in one go. With this approach, the code you will be receiving is very likely to be organized in sections/steps/stages or something similar. Machines are creative but also predictable, and this plays in our favor.

Action — Execute the plan one stage at a time, and check each result before moving to the next.

Action — If you think you solved the problem in a single attempt, read the box below. If you found some issue, mismatch, error message or anything that didn’t go as expected, go to Part 7 directly.

Before moving to Part 7

There is a chance that the solution AI has provided you on the first go is enough to solve your problem correctly. Just to make sure, let’s do a simple check: these are the final counts you should get for the exercise to be completed:

  reported_symptoms    n
1             fever 2396
2    neck stiffness  997
3         petechiae 1268
4          vomiting 1211
5              <NA>  122

Now we are really hoping this result caused you some confusion. Obviously, some decisions were made about the symptoms’ categories, that we didn’t prompt you about - on purpose. Regardless, the numbers should add up with the petechiae-rash and the many NA-like categories in the data. Also, the fever, neck stiffness and vomiting categories should match easily.

How did we make it there? Just read the next section for a step-by-step explanation

Part 7 · What you should have noticed

Real data doesn’t announce its problems. Some of them could be easy to spot on the go, because they cause error messages when they contradict AI’s code. You can only describe what you know, the rest requires interaction. When you want to transform a list of characters into single columns to test presence/absence, R doesn’t know that a single symptom may have multiple spellings (like when cleaning, remember?). Or that they used different names to refer to the same manifestation.

Here’s what was actually hiding in reported_symptoms, and why each one mattered.

Capitalisation and whitespace

Different clinicians typed this column by hand, some in lowercase, some capitalised, some with a stray leading or trailing space. To R, "fever" and " Fever" are two completely different strings. Left unresolved, this silently split a single symptom into two or three separate categories in your count.

We have functions that take care of that, but only if you realize you need to take care of it. Maybe you realized it when visually exploring the data at the beginning. But alternatively you should have asked for a first check the moment you splat the column and had the first possibility of counting categories.

If that happened to you, there is a simple fix which consist on asking AI for solutions to handle spelling mistakes and mismatches. Some of them just consist on erasing blank spaces or dealing with capitalization, but some other may require recoding of categories.

A symptom listed twice

A handful of cases had the same symptom appear twice within one string, a plausible double-entry error at the point of care or withing the linkage tool (the data scientist even warned you about some bugs…).

One strategy AI will have suggested you (most likely), is to split the many symptoms per row into one individual row per symptom and patient. From there, you would pivot_wider() to get the columns to count. If, lucky you, have added values_fill = 0 on the code, and duplicates are present, pivot_wider() will enter in conflict and drop an error

Again, this depends on the strategy AI produces for you. Maybe it already accounted for the possibility of duplicates and proposed you distinct() in the pipeline. Maybe you didn’t include the fill value argument, and the error did not appear.

Is rash the same as petechiae?

Unlike the two issues above, this one isn’t a formatting problem at all — it’s a judgement and clinical call.

Rash and petechiae are not the same. But they could refer to a very similar clinical manifestation, relevant in some severe acute presentations of IMD. You don’t know who wrote the notes, or interpreted them. If you are not a doctor or nurse, you may not realize about this issue. And you have to make decisions.

No function will resolve this for you, and no AI assistant should be trusted to make this call silently on your behalf. It requires domain knowledge you have and it doesn’t.

Deciding to treat rash as petechiae is a defensible clinical decision, not a cleaning step; deciding not to would be equally defensible, and would change your final counts. The important part is noticing the fork exists before your pivot decides it for you by creating two separate columns.

Two different kinds of “nothing”

Some cases simply had a blank cell — no symptom was ever recorded. Others had a clinician explicitly type "None", "N/A", or "Unknown" instead of leaving the field empty. These are not the same phenomenon: some are missing data, the others are a deliberate (if unhelpfully vague) entry.

For this exercise, you can make the decision of grouping all of them in a single NA or “Unknown” category. You could also differentiate missing from unknown values. The decision would affect the final count, and again requires domain knowledge.

For the proposed solution, we decided to group them all under NA, for simplicity. It isn’t a fifth symptom. It’s the visible trace of every case that had nothing to report, however that “nothing” was originally recorded.


Exercise summary

You didn’t just reshape a column today. You practiced a way of working that matters more than any single function: understand the problem yourself first, hand the AI a destination rather than a set of steps, treat its first answer as a proposal to test rather than a solution to trust, and when something breaks, come back with a sharper question instead of starting over. This loop — understand, define, propose, verify, refine — is the actual shape of AI-assisted programming, and it will serve you long after pivot_wider() syntax has faded from memory.

💡 Show solution — only after trying yourself!

(Instructor reference — not the only valid path; several strategies can reach the same destination)

This is an example prompt for this challenge:

“I’m a field epidemiologist working in R with the tidyverse. I have a dataframe called imd_s6, one row per case of invasive meningococcal disease. One column, reported_symptoms, holds a single string per case listing the symptoms a clinician observed, separated by semicolons (e.g. “fever;vomiting”). These were typed freehand by different clinicians, so I expect inconsistencies I haven’t fully catalogued — things like inconsistent capitalisation or spacing, and possibly other issues I haven’t spotted yet.

My goal: end up with one row per case, one column per distinct symptom, with a 1/0 indicator for whether that symptom was reported, so I can count how many cases reported each one.

Before giving me any code, walk me through your proposed strategy in plain language, step by step. Default to tidyverse functions, and explain the purpose of each step as you go rather than assuming I already know why it’s needed. I’ll test each step against my own data before asking you to continue to the next.”

# Step 1 — reshape from one string per case into one row per symptom mentioned
imd_long <- imd_s6 %>%
  mutate(reported_symptoms = str_trim(str_to_lower(reported_symptoms))) %>%
  separate_longer_delim(reported_symptoms, delim = ";") %>%
  mutate(reported_symptoms = str_trim(reported_symptoms))

# Checkpoint — every distinct value should now be either a real symptom, a
# resolvable synonym, or an explicit missing-data placeholder. Anything else
# needs another look before moving on.
imd_long %>% count(reported_symptoms)

# Step 2 — resolve the synonym ("rash") and the missing-data placeholders
# ("none" / "n/a" / "unknown") found at the checkpoint above. Cases with no
# real symptom become NA here — they are not dropped, only their value is.
imd_long <- imd_long %>%
  mutate(
    reported_symptoms = case_match(
      reported_symptoms,
      "rash" ~ "petechiae",
      c("none", "n/a", "unknown") ~ NA,
      .default = reported_symptoms
    ),
    present = 1
  ) %>%
  distinct()   # collapses the duplicated-token rows caught during exploration

# Step 3 — reshape to one row per case, one column per symptom.
# case_id travels along inside imd_long, so pivot_wider() uses it (along with
# every other untouched column) to anchor each row's identity correctly.
# A useful side effect: cases with no real symptom collapse into their own
# "NA" column instead of vanishing — see the checkpoint below.
imd_wide <- imd_long %>%
  pivot_wider(
    names_from = reported_symptoms,
    values_from = present,
    values_fill = 0
  )

# Checkpoint — row count must match the source data exactly. If it doesn't,
# some case disappeared silently somewhere upstream.
nrow(imd_wide) == nrow(imd_s6)

# Step 4 — count per symptom. The `NA` column is a bookkeeping artefact from
# Step 3 (cases with nothing to report), not a fifth symptom — kept here only
# to cross-check its count against the 122 confirmed empty/placeholder cases.
imd_wide %>%
  summarise(across(c(fever, `neck stiffness`, petechiae, vomiting, `NA`), sum))
The design assistant
Source Code
---
title: "The coding assistant"
subtitle: "Session 6 practical exercises"
---

```{r}
#| include: false
library(webexercises)
pacman::p_load(tidyverse, rio, here)

imd_s6 <- import(here("Data", "Clean", "IMD_Sample_Clean_S6E3.rds"))
```

An email at your inbox calls your attention. It comes from the Institute's data science team, right to you:

*"Sorry it took us more time than expected. Kassandra asked us to prepare more some data for you, now that we have our newest linkage to the clinical capture system to the surveillance registry in place. We have collected available clinical information about your cases, you'll find it as a new column, `reported_symptoms`, in the dataset I'm sending you from the last cleaned version we received. The tool is still ongoing fine-tuning and bug-fixing, and since it's based on free text entries from the clinicians, it may not be perfect tidy data yet. Good luck! "*

Kassandra was cc'ed, and soon enough you receive a second mail:

*"I want a symptom frequency count for the case definition review next week, send it to me before tomorrow so I have time to review it carefully during the weekend"*

Doesn't seem like a complicated petition, considering what you have already done so far... right?

**Action** — Open the new dataframe and explore the variable `reported_symptoms`. Can you produce that table right away?

## Part 6 · Understand and plan before prompting

This is a different kind of mess to anything you have cleaned so far in this course. It is not a typo, not an inconsistent label, not a wrong data type. It is a **structural** problem: the information you need exists, but it is not organised the way you need it to answer Kassandra's question.

You will use an AI assistant to help you solve this. But before you open that chat window, there is something worth sitting with for a moment:

- What do you need to accomplish, exactly?

- Which problems can you identify from the beginning?

- How will you move from what you have to where you need to go?

The skill this exercise is really about is not any single R function, it is how you work *with* an AI assistant through a problem like this: how you frame it, how you check its answers, and how you come back with a sharper question when something breaks. A vague question to an AI assistant about data you do not understand tends to produce a vague, generic answer — one that solves *a* version of the problem, not necessarily *yours*.

So before writing a single prompt, spend a few minutes with the column on your own terms:

**Action** — ***Understand your starting point***: explore `reported_symptoms` using tools you already know: `unique()`, `count()`, and manually inspecting a handful of raw rows. Take note of everything you notice: How many symptoms seem to appear per case? Is the order consistent? Are there any values that surprise you? This is your starting point.

**Action** — ***Define the destination before the path***: before asking for any code solution, think about how your data should look like once cleaned.

With this in mind, we can go back to our AI tool of choice.

### Ask for a strategy, not a solution

By combining the information from your starting point and your desired destination, the first prompt emerges organically: how do I arrive from A to B? This is a pathway, not a code solution. We will deliberately avoid asking for code now, that comes later.

**Action** — Start building your prompt, including:

- Contextual baselien information: who you are, what are you doing, how do you want AI to reply, with which libraries or approaches, in which tone. You should ask explicitly AI to use `Tidyverse`, and explain you the proposed solutions (functions, logic, etc.) by the way.

- Your specific problem: starting and destination points

- Your petition: a step-by-step strategy, in plain language first without code

**Action** — Check the proposed strategy against what you noticed when exploring your data. Does it account for what you saw? Would it handle a case with only one symptom? A case with none?

::: {.callout-tip collapse="true"}
## If the first suggestion sounds too simple

A very common first suggestion is to split the string into a fixed number of columns — `symptom_1`, `symptom_2`, and so on. Think back to what you noticed about ordering before accepting that. Does the *first* value in the string always mean the same thing across cases?
:::

**Action** — A little back-and-forth to refine the plan is usually needed. Check the plan, and prompt again if you are not convinced: identify the problematic bits and ask for correction, or explanations until the plan is set

The AI tool is only as good as your prompt, but it is also quite clever for spotting inconsistencies. Is the strategy it proposed seem truly odd and far from your original ideas, ask why, and you may be surprised.

### Execute and verify - step by step

This is the part where the plan meets the actual data.

**Action** — Now ask AI for the code, once the plan is agreed and fixed.

**Do not run the whole strategy in one go**. With this approach, the code you will be receiving is very likely to be organized in sections/steps/stages or something similar. Machines are creative but also predictable, and this plays in our favor.

**Action** — Execute the plan **one stage at a time**, and **check** each result before moving to the next.

**Action** — If you think you solved the problem in a single attempt, read the box below. If you found some issue, mismatch, error message or anything that didn't go as expected, go to Part 7 directly.

::: callout-important
## Before moving to Part 7

There is a chance that the solution AI has provided you on the first go is enough to solve your problem correctly. Just to make sure, let's do a simple check: these are the final counts you should get for the exercise to be completed:

``` r
  reported_symptoms    n
1             fever 2396
2    neck stiffness  997
3         petechiae 1268
4          vomiting 1211
5              <NA>  122
```

Now we are really hoping this result caused you some confusion. Obviously, some decisions were made about the symptoms' categories, that we didn't prompt you about - on purpose. Regardless, the numbers should add up with the `petechiae-rash` and the many `NA`-like categories in the data. Also, the `fever`, `neck stiffness` and `vomiting` categories should match easily.

How did we make it there? Just read the next section for a step-by-step explanation
:::

## Part 7 · What you should have noticed

Real data doesn't announce its problems. Some of them could be easy to spot on the go, because they cause error messages when they contradict AI's code. You can only describe what you know, the rest requires interaction. When you want to transform a list of characters into single columns to test presence/absence, R doesn't know that a single symptom may have multiple spellings (like when cleaning, remember?). Or that they used different names to refer to the same manifestation.

Here's what was actually hiding in `reported_symptoms`, and why each one mattered.

### Capitalisation and whitespace

Different clinicians typed this column by hand, some in lowercase, some capitalised, some with a stray leading or trailing space. To R, `"fever"` and `" Fever"` are two completely different strings. Left unresolved, this silently split a single symptom into two or three separate categories in your count.

We have functions that take care of that, but only if you realize you need to take care of it. Maybe you realized it when visually exploring the data at the beginning. But alternatively you should have asked for a first check the moment you splat the column and had the first possibility of counting categories.

If that happened to you, there is a simple fix which consist on asking AI for solutions to handle spelling mistakes and mismatches. Some of them just consist on erasing blank spaces or dealing with capitalization, but some other may require recoding of categories.

### A symptom listed twice

A handful of cases had the same symptom appear twice within one string, a plausible double-entry error at the point of care or withing the linkage tool (the data scientist even warned you about some bugs...).

One strategy AI will have suggested you (most likely), is to split the many symptoms per row into one individual row per symptom and patient. From there, you would `pivot_wider()` to get the columns to count. If, lucky you, have added `values_fill = 0` on the code, and duplicates are present, `pivot_wider()` will enter in conflict and drop an error

Again, this depends on the strategy AI produces for you. Maybe it already accounted for the possibility of duplicates and proposed you `distinct()` in the pipeline. Maybe you didn't include the fill value argument, and the error did not appear.

### Is `rash` the same as `petechiae`?

Unlike the two issues above, this one isn't a formatting problem at all — it's a judgement and clinical call.

Rash and petechiae are not the same. But they could refer to a very similar clinical manifestation, relevant in some severe acute presentations of IMD. You don't know who wrote the notes, or interpreted them. If you are not a doctor or nurse, you may not realize about this issue. And you have to make decisions.

No function will resolve this for you, and no AI assistant should be trusted to make this call silently on your behalf. It requires domain knowledge you have and it doesn't.

Deciding to treat `rash` as `petechiae` is a defensible clinical decision, not a cleaning step; deciding *not* to would be equally defensible, and would change your final counts. The important part is noticing the fork exists before your pivot decides it for you by creating two separate columns.

### Two different kinds of "nothing"

Some cases simply had a blank cell — no symptom was ever recorded. Others had a clinician explicitly type `"None"`, `"N/A"`, or `"Unknown"` instead of leaving the field empty. These are not the same phenomenon: some are missing data, the others are a deliberate (if unhelpfully vague) entry.

For this exercise, you can make the decision of grouping all of them in a single `NA` or "Unknown" category. You could also differentiate missing from unknown values. The decision would affect the final count, and again requires domain knowledge.

For the proposed solution, we decided to group them all under `NA`, for simplicity. It isn't a fifth symptom. It's the visible trace of every case that had nothing to report, however that "nothing" was originally recorded.

------------------------------------------------------------------------

## Exercise summary

You didn't just reshape a column today. You practiced a way of working that matters more than any single function: understand the problem yourself first, hand the AI a destination rather than a set of steps, treat its first answer as a proposal to test rather than a solution to trust, and when something breaks, come back with a sharper question instead of starting over. This loop — understand, define, propose, verify, refine — is the actual shape of AI-assisted programming, and it will serve you long after `pivot_wider()` syntax has faded from memory.

::: {.callout-tip collapse="true"}
## 💡 Show solution — only after trying yourself!

*(Instructor reference — not the only valid path; several strategies can reach the same destination)*

This is an example prompt for this challenge:

> "*I'm a field epidemiologist working in R with the tidyverse. I have a dataframe called imd_s6, one row per case of invasive meningococcal disease. One column, reported_symptoms, holds a single string per case listing the symptoms a clinician observed, separated by semicolons (e.g. "fever;vomiting"). These were typed freehand by different clinicians, so I expect inconsistencies I haven't fully catalogued — things like inconsistent capitalisation or spacing, and possibly other issues I haven't spotted yet.*
>
> *My goal: end up with one row per case, one column per distinct symptom, with a 1/0 indicator for whether that symptom was reported, so I can count how many cases reported each one.*
>
> *Before giving me any code, walk me through your proposed strategy in plain language, step by step. Default to tidyverse functions, and explain the purpose of each step as you go rather than assuming I already know why it's needed. I'll test each step against my own data before asking you to continue to the next.*"

``` r
# Step 1 — reshape from one string per case into one row per symptom mentioned
imd_long <- imd_s6 %>%
  mutate(reported_symptoms = str_trim(str_to_lower(reported_symptoms))) %>%
  separate_longer_delim(reported_symptoms, delim = ";") %>%
  mutate(reported_symptoms = str_trim(reported_symptoms))

# Checkpoint — every distinct value should now be either a real symptom, a
# resolvable synonym, or an explicit missing-data placeholder. Anything else
# needs another look before moving on.
imd_long %>% count(reported_symptoms)

# Step 2 — resolve the synonym ("rash") and the missing-data placeholders
# ("none" / "n/a" / "unknown") found at the checkpoint above. Cases with no
# real symptom become NA here — they are not dropped, only their value is.
imd_long <- imd_long %>%
  mutate(
    reported_symptoms = case_match(
      reported_symptoms,
      "rash" ~ "petechiae",
      c("none", "n/a", "unknown") ~ NA,
      .default = reported_symptoms
    ),
    present = 1
  ) %>%
  distinct()   # collapses the duplicated-token rows caught during exploration

# Step 3 — reshape to one row per case, one column per symptom.
# case_id travels along inside imd_long, so pivot_wider() uses it (along with
# every other untouched column) to anchor each row's identity correctly.
# A useful side effect: cases with no real symptom collapse into their own
# "NA" column instead of vanishing — see the checkpoint below.
imd_wide <- imd_long %>%
  pivot_wider(
    names_from = reported_symptoms,
    values_from = present,
    values_fill = 0
  )

# Checkpoint — row count must match the source data exactly. If it doesn't,
# some case disappeared silently somewhere upstream.
nrow(imd_wide) == nrow(imd_s6)

# Step 4 — count per symptom. The `NA` column is a bookkeeping artefact from
# Step 3 (cases with nothing to report), not a fifth symptom — kept here only
# to cross-check its count against the 122 confirmed empty/placeholder cases.
imd_wide %>%
  summarise(across(c(fever, `neck stiffness`, petechiae, vomiting, `NA`), sum))
```
:::

```{=html}
<script>
document.addEventListener("DOMContentLoaded", function() {
  var radiogroups = document.getElementsByClassName("webex-radiogroup");
  for (var i = 0; i < radiogroups.length; i++) {
    radiogroups[i].onchange = radiogroups_func;
  }
});
</script>
```

© 2026 – Intro to R Course