Skip to main content

Using Java Regex Functions on CFML Strings

I ran into this today while working on ColdTonica, and since it's something I'm still surprised people forget (including myself) I thought I'd share.

ColdTonica is a CFML clone of StatusNet (formerly Laconica), which is an open source PHP-based microblogging service similar (although vastly superior) to Twitter. As you might imagine, those simple 140-character notices you spend way too much of your day posting go through a lot of transformations before reaching the final form in which they are displayed, because the notices need to be parsed and manipulated to add things like links to tags, links to @ replies, shortening URLs, and so on. Honestly when I started studying the StatusNet code and saw what all goes on behind the scenes for such a seemingly simple service, I have to admit I was a bit surprised.

All the text manipulation of course involves a lot of regular expressions, and since for ColdTonica we're porting the PHP code over to CFML, it saves us a ton of time since all the regular expressions have already been written. Unfortunately there are some syntax issues with the regular expressions that render them incompatible with CFML and even Java, so it did take a bit of research and help from a friend of mine to start to unravel and convert them.

One of the issues I ran into while moving the PHP regular expressions over to CFML is that CFML doesn't have Unicode support in regular expressions (some nice info about Unicode in regex here, although the information about Java is dated), at least not without first converting Unicode to ASCII values and wrapping them all in Chr(). This is what I've discovered while messing with this at least; if this isn't correct I'm happy to be proven wrong.

Since the PHP regular expressions use the Perl syntax of x{hex_value_here}, which CFML doesn't support, converting the regular expressions was getting a bit messy. Java, however, does support the x syntax (though it didn't used to), but with a slightly different syntax. You can read more about Java regex syntax in Java 6 here.

During the course of this I was reminded of the fact that under the hood, CFML strings are Java strings, which means that rather than using functions like REReplaceNoCase() in CFML and converting the hex codes into something usable by Chr(), I can simply use Java's replaceAll() function on the String class. This lets me keep the PHP syntax more intact and do a lot less conversion research.

So the original PHP looks like this:

$r = preg_replace('/[x{0}-x{8}x{b}-x{c}x{e}-x{19}]/', '', $r);



And the CFML version using replaceAll() on the String class looks like this:

r.replaceAll("/[x00-x08x0B-x0Cx0E-x19]/", "");



At least I think that's right. ;-) I still need to test all of this out, but as I convert the rest of these it'll be much simpler to go this route and keep things in hex as opposed to converting everything to CFML-compatible Unicode regex syntax.

The moral of the story is you can do a lot in CFML by leveraging the underlying Java functionality, and this doesn't apply only to the String class. So if you run into things that are a bit weird to try and accomplish in CFML check the Java docs and see what additional functionality you have available. You'll probably be surprised at what you learn!

Comments

Popular posts from this blog

Installing and Configuring NextPVR as a Replacement for Windows Media Center

If you follow me on Google+ you'll know I had a recent rant about Windows Media Center, which after running fine for about a year suddenly decided as of January 29 it was done downloading the program guide and by extension was therefore done recording any TV shows.

I'll spare you more ranting and simply say that none of the suggestions I got (which I appreciate!) worked, and rather than spending more time figuring out why, I decided to try something different.

NextPVR is an awesome free (as in beer, not as in freedom unfortunately ...) PVR application for Windows that with a little bit of tweaking handily replaced Windows Media Center. It can even download guide data, which is apparently something WMC no longer feels like doing.

Background I wound up going down this road in a rather circuitous way. My initial goal for the weekend project was to get Raspbmc running on one of my Raspberry Pis. The latest version of XBMC has PVR functionality so I was anxious to try that out as a …

Setting Up Django On a Raspberry Pi

This past weekend I finally got a chance to set up one of my two Raspberry Pis to use as a Django server so I thought I'd share the steps I went through both to save someone else attempting to do this some time as well as get any feedback in case there are different/better ways to do any of this.

I'm running this from my house (URL forthcoming once I get the real Django app finalized and put on the Raspberry Pi) using dyndns.org. I don't cover that aspect of things in this post but I'm happy to write that up as well if people are interested.

General Comments and Assumptions

Using latest Raspbian “wheezy” distro as of 1/19/2013 (http://www.raspberrypi.org/downloads)We’lll be using Nginx (http://nginx.org) as the web server/proxy and Gunicorn (http://gunicorn.org) as the WSGI serverI used http://www.apreche.net/complete-single-server-django-stack-tutorial/ heavily as I was creating this, so many thanks to the author of that tutorial. If you’re looking for more details on …

The Definitive Guide to CouchDB Authentication and Security

With a bold title like that I suppose I should clarify a bit. I finally got frustrated enough with all the disparate and seemingly incomplete information on this topic to want to gather everything I know about this topic into a single place, both so I have it for my own reference but also in the hopes that it will help others.Since CouchDB is just an HTTP resource and can be secured at that level along the same lines as you'd secure any HTTP resource, I should also point out that I will not be covering things like putting a proxy in front of CouchDB, using SSL with CouchDB, or anything along those lines. This post is strictly limited to how authentication and security work within CouchDB itself.CouchDB security is powerful and granular but frankly it's also a bit quirky and counterintuitive. What I'm outlining here is my understanding of all of this after taking several runs at it, reading everything I could find on the Internet (yes, the whole Internet!), and a great deal…