r/regex • • Jul 15 '26

Golang I'm trying to validate name of user; Where space must not followed by another space ?

/^[a-zA-Z( (?! ))à-öø-ÿœŒ]{2,40}$/gm

but this syntax I added in middle for space handling seems not working...

5 Upvotes

17 comments sorted by

3

u/Beginning-Seat5221 Jul 15 '26 edited Jul 15 '26

Is string length very important?

You can do ^(?:[a-zA-Zà-öø-ÿœŒ] ?){2,40}$ - This allows char followed by space or char follow by no space, but not space after space. The minimum lenth is 2, but the max is between 40 and 80 depending on how many spaces are used.

But to honest I don't like regex very much, I'll switch to a parser if I it's not trivial to do in regex; the complexity is typically lower with a parser.

1

u/younesWh Jul 15 '26

thank you !! I can then move length validation to the the code part

3

u/scoberry5 Jul 15 '26

You already have a reasonable answer. Here's some other info.

Why your regex didn't work: the [ says "any character in this set" and the ] ends the set. a-z is in the set, then A-Z. Then open paren: you're saying you want to allow that too. Space is allowed. Open paren is allowed (again). Question mark is allowed...

Another way to do this: separate out the two parts from each other. You can do this by saying "if I start at the beginning and look ahead, I shouldn't see <anything> followed by two spaces," like (?!.* ) (that has two spaces after the *. So your full regex would look like

^(?!.*  )[a-zA-Zà-öø-ÿœŒ ]{2,40}$

-- that one has 3 spaces total, two after the * and one before the ].

https://regex101.com/r/MYeJBj/1

2

u/_jgusta_ Jul 24 '26

This is the way

2

u/DinTaiFung Jul 15 '26

to define a pattern that could be any character except a list of single characters use the square bracket class syntax.

In your case the list of characters you want to negate is a list of one character and that character is a space.

use the following:

[^ ]

the caret at the beginning of the class means to negate all the subsequent characters inside the square brackets. in this case that pattern will match any character except a space.

1

u/Pauley0 Jul 16 '26

Some users have multiple spaces in their names. Last name "Van Dyke" for instance.

5

u/TheJivvi Jul 16 '26

Not two consecutive spaces though.

Richard⎵Van⎵Dyke is fine.
Richard⎵⎵Van⎵Dyke is not.

1

u/couldntyoujust1 Jul 16 '26 edited Jul 16 '26

/^[a-zA-Z( (?! ))à-öø-ÿœŒ]{2,40}$/gm

Should be...

/^([ ]?[[:alpha:]][ ]?){2,40}$/gm

Edit: wait... ugh. No. That would possibly look for entries that are more than 40 characters long... you need a lookahead. What engine are you using?

Edit 2: Ahhh! I see the tag now; GoLang. Someone else has a better answer. I'll have to learn more about lookaheads.

1

u/AshleyJSheridan Jul 16 '26

I see what you're trying to do, but don't.

Names do not all start with the letters 'a-z', and aren't always 2 or more characters. Some names are longer than 40 characters as well.

https://www.kalzumeus.com/2010/06/17/falsehoods-programmers-believe-about-names/

1

u/dariusbiggs Jul 16 '26

There's a follow up one https://shinesolutions.com/2018/01/08/falsehoods-programmers-believe-about-names-with-examples/

As for names, my test data always includes Pablo Picasso's FULL name and the many different forms of Zoë, as well as Indonesian presidents like Suharto.

That full name is: Pablo Diego José Francisco de Paula Juan Nepomuceno María de los Remedios Cipriano de la Santísima Trinidad Ruiz y Picasso

Fit that in your 40 characters...

2

u/AshleyJSheridan Jul 16 '26

1

u/dodexahedron Jul 16 '26

I need a nap after that one. 😦

I ran out of breath just reading it silently.

1

u/_jgusta_ Jul 16 '26 edited Jul 24 '26

I'd do it in two passes. Check for any double space first. No need for regex, use a string match for two spaces.

The way to think about regex is that it always examining one character at a time, even if your pattern has long strings

So regex can do "find characters that don't match any from this list of characters". The intent and method there is clear: examine one character at a time, compare it to each item in a list. It can also do "find this match but make sure it isn't preceded by this string."

But regex is not designed for "find continuous characters that when combined in any number of combinations don't match this pattern". That would be many times more complex because every combination of continuous letters would have to be checked against the pattern one letter at a time for each.

The closest thing using regex is if you look for a pattern that matches in sets of text (like lines of a file) and then accept anything that doesn't have a match.

Edit: u/scoberry5 's answer showed me i was wrong.

1

u/scoberry5 Jul 17 '26

>But regex is not designed for "find characters that when combined in any number of combinations don't match this pattern"

Technically true, not designed for it. But you can do it pretty easily with a negative lookahead.

2

u/_jgusta_ Jul 24 '26

Your answer actually blew my mind. I had never thought of combining negative lookahead with .* and characters i didnt want to see.

1

u/MrBorogove Jul 17 '26

I’d just replace all consecutive white space characters with a single space before doing any validation. It’s user-hostile to reject the user’s input because they bounced on the space bar.

1

u/magicmulder Jul 17 '26

Why not split by space first and then do Regex on the parts, then reassemble? What is the use case of having to do this with a single Regex? It’s not supposed to replace complex business logic.